⌘K
← Can It Run?

Can the GeForce RTX 5090 run Mistral Small 3.2 24B?

Mistral Small 3.2 24B runs locally on any 24 GB GPU at Q4_K_M with room to spare: the model needs about 16.6 GB of VRAM with a 4k-token context. A 20 GB Radeon RX 7900 XT also fits at Q4_K_M, while 16 GB cards must drop to a lower quant or offload layers.

Fits comfortably

At Q4_K_M quantization with 4,096 context, the Mistral Small 3.2 24B requires approximately 16.64GB of VRAM. The GeForce RTX 5090 has 32GB available (15.36GB headroom).

VRAM Breakdown

ComponentSize (GB)Source
Model weights (Q4_K_M)14.52calculated
Loading overhead (10%)1.45calculated
KV cache @ 4,096 context0.67calculated
Total required16.64
GPU memory available32vendor_spec
Headroom / shortfall+15.36

Quantization Options

QuantWeights (GB)Total @ 4K (GB)Total @ 32K (GB)Fits GeForce RTX 5090?
Q4_K_M 14.52 16.64 21.34 Yes
Q5_K_M 17.24 19.63 24.33 Yes
Q6_K 19.97 22.64 27.34 Yes
Q8_0 25.71 28.95 33.65 Yes
FP16 48.4 53.91 58.61 No

Get the GeForce RTX 5090

Check current price on Amazon

Affiliate link — we earn a commission on qualifying purchases at no cost to you.

If It Doesn't Fit: Alternatives

Same GPU, Other Models

Frequently Asked Questions

Can the Mistral Small 3.2 24B run on 32GB VRAM?

Yes. At Q4_K_M quantization with 4K context, the Mistral Small 3.2 24B needs approximately 16.64GB. Your 32GB of usable memory has 15.36GB headroom.

What is the maximum context length I can use on the GeForce RTX 5090?

The Mistral Small 3.2 24B supports up to 131,072 tokens of context. Larger context increases KV-cache memory requirements. At 4K context the model needs ~16.64GB; at 32K it needs more.

Which quantization should I use with the GeForce RTX 5090?

Q4_K_M is the recommended starting point — it offers the best speed-to-quality tradeoff. If you have VRAM headroom, Q6_K or Q8_0 improves quality at the cost of speed. FP16 is only for inference servers with ample memory.

Can a 16 GB GPU run Mistral Small 3.2 24B?

Not at Q4_K_M with a 4k context, which needs about 16.6 GB of VRAM. A 16 GB GPU runs the model at Q4_K_S or IQ4_XS with a shorter context, or offloads a few layers to system RAM at a speed cost.

How much VRAM does Mistral Small 3 24B need?

Mistral Small 3.2 24B needs about 16.6 GB of VRAM at Q4_K_M with a 4k context: 14.52 GB of weights plus 10% overhead plus a 0.67 GB KV cache.

Related Pages

Full Mistral Small 3.2 24B fit page · All AI models · GeForce RTX 5090 specs · VRAM calculator

VRAM figures are calculated from model architecture parameters using established formulas. Benchmark data is measured and cited per row. GPU memory is vendor_spec from manufacturer datasheets.

Methodology · Editorial policy · Affiliate disclosure