Can you run Mistral Medium 3.5 128B locally?
Mistral Medium 3.5 family · 2026 · 127.7B parameters (127.7B activated per token) · Hugging Face model card
Mistral Medium 3.5 128B is a dense model with 127.7B total open weights (official card: Dense 128B parameters; HF safetensors.total=127,704,210,176). Native context is 262,144 tokens (card: 256k context window). License is other (Mistral Research License per repository). Local VRAM fit at Q4-class is roughly 96 GB (calculated) multi-GPU server territory. Measured speeds: not yet published here.
Minimum: 96 GB+ memory (Q4-class, calculated) · Recommended: 96 GB+ for comfortable context headroom (Q4-class, calculated)
How much VRAM does Mistral Medium 3.5 128B need at each quantization?
Mistral Medium 3.5 128B needs 85.8 GB of VRAM at Q4_K_M. The table below lists weights-only size and total VRAM including overhead for each common quantization level with 4k and 32k token contexts.
| Quantization | Bits / weight | Weights only | Total + KV @4k ctx | Total + KV @32k ctx |
|---|---|---|---|---|
| Q4_K_M | 4.8 | 76.6 GB | 85.8 GB | 96.1 GB |
| Q5_K_M | 5.7 | 91.0 GB | 101.6 GB | 111.9 GB |
| Q6_K | 6.6 | 105.4 GB | 117.4 GB | 127.7 GB |
| Q8_0 | 8.5 | 135.7 GB | 150.7 GB | 161.1 GB |
| FP16 | 16 | 255.4 GB | 282.4 GB | 292.8 GB |
Method: weights = parameters × bits-per-weight (Q4_K_M ≈ 4.8, Q5_K_M ≈ 5.7, Q6_K ≈ 6.6, Q8_0 ≈ 8.5, FP16 = 16), plus 10% loading overhead, plus KV cache from the verified architecture config (88 layers, 8 KV heads, 128 head dim). Source: HF-verified 2026-09-28 via api (gated=false private=false params=127.7B safetensors.total=127704210176 downloadable_safetensors_sum=267212267176 ctx=262144 license=other repo=mistralai/Mistral-Medium-3.5-128B@22b2b868a15677cfa6061277ed2f653d1349a9ab layers=88 kv_heads=8). MoE: num_experts=None num_experts_per_tok=None active_params_b=127.7 (vendor card; VRAM fit uses active, not dense-equivalent). VRAM figures are documented calculations, not measurements. No tok/s invented.
Which GPUs can run Mistral Medium 3.5 128B locally?
At Q4_K_M with a 4k context, Mistral Medium 3.5 128B needs 85.8 GB of VRAM. The lists below are computed live from our GPU database and grouped by how much headroom the card has. Disclosure: CompareAIHardware.com participates in the Amazon Associates program and earns from qualifying purchases through links on this page. Affiliate relationships do not influence our recommendations.
Runs Mistral Medium 3.5 128B comfortably (25%+ VRAM headroom)
| GPU | VRAM | Bandwidth | Type | |
|---|---|---|---|---|
| H200 SXM | 141 GB | 4800 GB/s | Data Center GPU | |
| AMD Instinct MI300X | 192 GB | 5300 GB/s | Data Center GPU | |
| NVIDIA B200 | 192 GB | 8000 GB/s | Data Center GPU | |
| AMD Instinct MI325X | 256 GB | 6000 GB/s | Data Center GPU | |
| AMD Instinct MI355X | 288 GB | 8000 GB/s | Data Center GPU | |
| NVIDIA B300 (Blackwell Ultra) | 288 GB | 8000 GB/s | Data Center GPU |
Minimum GPUs that fit Mistral Medium 3.5 128B at Q4_K_M
These GPUs hold the model but leave little headroom — keep contexts short.
| GPU | VRAM | Bandwidth | Type | |
|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell | 96 GB | 1792 GB/s | Pro GPU | Check price |
Needs 2+ GPUs to run Mistral Medium 3.5 128B
One of these cards is too small on its own, but a pair (tensor or pipeline parallel, ~90% efficiency) covers the 85.8 GB requirement. See our multi-GPU guide for setup.
| GPU | VRAM | Bandwidth | Type | |
|---|---|---|---|---|
| AMD Radeon Pro W7900 | 48 GB ×2 | 864 GB/s | Pro GPU | Check price |
| NVIDIA RTX PRO 5000 Blackwell | 48 GB ×2 | 1344 GB/s | Pro GPU | |
| RTX 6000 Ada | 48 GB ×2 | 960 GB/s | Pro GPU | Check price |
| RTX A6000 | 48 GB ×2 | 768 GB/s | Pro GPU | Check price |
| H100 SXM | 80 GB ×2 | 3350 GB/s | Data Center GPU | |
| NVIDIA A100 80GB SXM | 80 GB ×2 | 2039 GB/s | Data Center GPU | |
| NVIDIA H100 PCIe 80GB | 80 GB ×2 | 2039 GB/s | Data Center GPU |
Can a Mac run Mistral Medium 3.5 128B?
Yes — these Apple Silicon machines fit Mistral Medium 3.5 128B at Q4_K_M, since macOS lets the GPU use about 75% of unified memory. Generation speed is bound by memory bandwidth, so the GB/s column matters as much as capacity.
| Apple Silicon | Unified memory | Usable by GPU (~75%) | Bandwidth |
|---|---|---|---|
| Apple M3 Max | 128 GB | 96 GB | 400 GB/s |
| M4 Max (MacBook Pro) | 128 GB | 96 GB | 546 GB/s |
| M3 Ultra (Mac Studio, 256 GB) | 256 GB | 192 GB | 819 GB/s |
Can Mistral Medium 3.5 128B run on mini PCs, Jetson, or NPU devices?
These edge and NPU devices from our database have enough memory for Mistral Medium 3.5 128B at Q4_K_M. Their memory bandwidth is far below discrete GPUs, so expect a fraction of desktop generation speed.
| Device | Memory | Bandwidth | Type |
|---|---|---|---|
| AMD Ryzen AI 9 HX 370 (Strix Point) | 96 GB | 120 GB/s | NPU Chip |
Which pre-built systems can run Mistral Medium 3.5 128B?
These mini PCs and workstations from our database fit Mistral Medium 3.5 128B at Q4_K_M. Usable-memory figures are conservative: Windows shares about half of system RAM with the GPU by default (Linux can expose more), macOS lets Apple Silicon GPUs use about 75% of unified memory, and Linux unified-memory systems such as GB10 expose roughly 90%. Generation speed is bound by memory bandwidth, so compare the GB/s column before buying.
| System | Type | Memory | Usable for AI | Bandwidth | GPUs |
|---|---|---|---|---|---|
| ACEMAGIC F9A (Ryzen AI Max+ PRO 495) | Mini PC | 192 GB | 96 GB | — | — |
| BOXX APEXX 8R (1x RTX PRO 6000 Blackwell) | Workstation | 96 GB | 96 GB | 1792 GB/s | 1 |
| Chuwi UniBox AI495 Pro (Ryzen AI Max+ PRO 495) | Mini PC | 192 GB | 96 GB | 273 GB/s | — |
| GEEKOM A9 Mega AI Workstation | Workstation | 128 GB | 96 GB | — | 1 |
| GMKtec EVO-X5 Pro (Ryzen AI Max+ PRO 495) | Mini PC | 192 GB | 96 GB | — | — |
| NOVATECH RTX PRO 6000 AI Workstation | Workstation | 96 GB | 96 GB | — | 1 |
| System76 Thelio Major (1x RTX PRO 6000 Blackwell) | Workstation | 96 GB | 96 GB | 1792 GB/s | 1 |
| System76 Thelio Major (2x RTX 6000 Ada) | Workstation | 96 GB | 96 GB | 960 GB/s | 2 |
| Lenovo ThinkStation PGX | Workstation | 128 GB | 115.2 GB | 273 GB/s | 1 |
| NVIDIA DGX Spark | Workstation | 128 GB | 115.2 GB | 273 GB/s | 1 |
| BIZON G3000 G2 (4x RTX 5090) | Workstation | 128 GB | 128 GB | 1792 GB/s | 4 |
| Apple Mac Studio M2 Ultra (192GB) | Workstation | 192 GB | 144 GB | 800 GB/s | 1 |
| HP Z8 Fury G5 (4x RTX 6000 Ada) | Workstation | 192 GB | 192 GB | 960 GB/s | 4 |