Can the GeForce RTX 5090 run Mistral Small 3.2 24B?
Mistral Small 3.2 24B runs locally on any 24 GB GPU at Q4_K_M with room to spare: the model needs about 16.6 GB of VRAM with a 4k-token context. A 20 GB Radeon RX 7900 XT also fits at Q4_K_M, while 16 GB cards must drop to a lower quant or offload layers.
At Q4_K_M quantization with 4,096 context, the Mistral Small 3.2 24B requires approximately 16.64GB of VRAM. The GeForce RTX 5090 has 32GB available (15.36GB headroom).
VRAM Breakdown
| Component | Size (GB) | Source |
|---|---|---|
| Model weights (Q4_K_M) | 14.52 | calculated |
| Loading overhead (10%) | 1.45 | calculated |
| KV cache @ 4,096 context | 0.67 | calculated |
| Total required | 16.64 | |
| GPU memory available | 32 | vendor_spec |
| Headroom / shortfall | +15.36 |
Quantization Options
| Quant | Weights (GB) | Total @ 4K (GB) | Total @ 32K (GB) | Fits GeForce RTX 5090? |
|---|---|---|---|---|
| Q4_K_M | 14.52 | 16.64 | 21.34 | Yes |
| Q5_K_M | 17.24 | 19.63 | 24.33 | Yes |
| Q6_K | 19.97 | 22.64 | 27.34 | Yes |
| Q8_0 | 25.71 | 28.95 | 33.65 | Yes |
| FP16 | 48.4 | 53.91 | 58.61 | No |
Get the GeForce RTX 5090
Affiliate link — we earn a commission on qualifying purchases at no cost to you.
If It Doesn't Fit: Alternatives
Same GPU, Other Models
Frequently Asked Questions
Can the Mistral Small 3.2 24B run on 32GB VRAM?
Yes. At Q4_K_M quantization with 4K context, the Mistral Small 3.2 24B needs approximately 16.64GB. Your 32GB of usable memory has 15.36GB headroom.
What is the maximum context length I can use on the GeForce RTX 5090?
The Mistral Small 3.2 24B supports up to 131,072 tokens of context. Larger context increases KV-cache memory requirements. At 4K context the model needs ~16.64GB; at 32K it needs more.
Which quantization should I use with the GeForce RTX 5090?
Q4_K_M is the recommended starting point — it offers the best speed-to-quality tradeoff. If you have VRAM headroom, Q6_K or Q8_0 improves quality at the cost of speed. FP16 is only for inference servers with ample memory.
Can a 16 GB GPU run Mistral Small 3.2 24B?
Not at Q4_K_M with a 4k context, which needs about 16.6 GB of VRAM. A 16 GB GPU runs the model at Q4_K_S or IQ4_XS with a shorter context, or offloads a few layers to system RAM at a speed cost.
How much VRAM does Mistral Small 3 24B need?
Mistral Small 3.2 24B needs about 16.6 GB of VRAM at Q4_K_M with a 4k context: 14.52 GB of weights plus 10% overhead plus a 0.67 GB KV cache.
Related Pages
Full Mistral Small 3.2 24B fit page · All AI models · GeForce RTX 5090 specs · VRAM calculator
VRAM figures are calculated from model architecture parameters using established formulas. Benchmark data is measured and cited per row. GPU memory is vendor_spec from manufacturer datasheets.