⌘K
← All AI Models

Can you run Qwen3 235B A22B locally?

Qwen3 family · · 235.1B parameters (22B activated per token) · Hugging Face model card

Qwen3 235B A22B is a mixture-of-experts model with 22.0B active parameters per token and 235.1B total open-weight model with a 40,960-token context window, usable locally with roughly 192 GB of memory at Q4-class quantization (calculated; measured speeds vary by hardware).

Minimum: 192 GB+ memory (Q4_K_M, calculated)  ·  Recommended: 192 GB+ memory for full context (Q4_K_M, calculated)

Parameters (total)
235.1B
Activated per token
22B (MoE)
Context window
40,960 tokens
Architecture source
HF config.json (verified)

How much VRAM does Qwen3 235B A22B need at each quantization?

Qwen3 235B A22B needs 156.0 GB of VRAM at Q4_K_M. The table below lists weights-only size and total VRAM including overhead for each common quantization level with 4k and 32k token contexts.

QuantizationBits / weightWeights onlyTotal + KV @4k ctxTotal + KV @32k ctx
Q4_K_M 4.8 141.1 GB 156.0 GB 161.5 GB
Q5_K_M 5.7 167.5 GB 185.0 GB 190.6 GB
Q6_K 6.6 194.0 GB 214.1 GB 219.7 GB
Q8_0 8.5 249.8 GB 275.6 GB 281.1 GB
FP16 16 470.2 GB 518.0 GB 523.5 GB

Method: weights = parameters × bits-per-weight (Q4_K_M ≈ 4.8, Q5_K_M ≈ 5.7, Q6_K ≈ 6.6, Q8_0 ≈ 8.5, FP16 = 16), plus 10% loading overhead, plus KV cache from the verified architecture config (94 layers, 4 KV heads, 128 head dim). Source: HF-verified 2026-09-12 via api (params=235.09B ctx=40960 license=apache-2.0); VRAM figures are documented calculations, not measurements

Which GPUs can run Qwen3 235B A22B locally?

At Q4_K_M with a 4k context, Qwen3 235B A22B needs 156.0 GB of VRAM. The lists below are computed live from our GPU database and grouped by how much headroom the card has. Disclosure: CompareAIHardware.com participates in the Amazon Associates program and earns from qualifying purchases through links on this page. Affiliate relationships do not influence our recommendations.

Runs Qwen3 235B A22B comfortably (25%+ VRAM headroom)

GPUVRAMBandwidthType
AMD Instinct MI325X 256 GB 6000 GB/s Data Center GPU
AMD Instinct MI355X 288 GB 8000 GB/s Data Center GPU
NVIDIA B300 (Blackwell Ultra) 288 GB 8000 GB/s Data Center GPU

Minimum GPUs that fit Qwen3 235B A22B at Q4_K_M

These GPUs hold the model but leave little headroom — keep contexts short.

GPUVRAMBandwidthType
AMD Instinct MI300X 192 GB 5300 GB/s Data Center GPU
NVIDIA B200 192 GB 8000 GB/s Data Center GPU

Needs 2+ GPUs to run Qwen3 235B A22B

One of these cards is too small on its own, but a pair (tensor or pipeline parallel, ~90% efficiency) covers the 156.0 GB requirement. See our multi-GPU guide for setup.

GPUVRAMBandwidthType
NVIDIA RTX PRO 6000 Blackwell 96 GB ×2 1792 GB/s Pro GPU Check price
H200 SXM 141 GB ×2 4800 GB/s Data Center GPU

Can a Mac run Qwen3 235B A22B?

Yes — these Apple Silicon machines fit Qwen3 235B A22B at Q4_K_M, since macOS lets the GPU use about 75% of unified memory. Generation speed is bound by memory bandwidth, so the GB/s column matters as much as capacity.

Apple SiliconUnified memoryUsable by GPU (~75%)Bandwidth
M3 Ultra (Mac Studio) 512 GB 384 GB 819 GB/s

Which pre-built systems can run Qwen3 235B A22B?

These mini PCs and workstations from our database fit Qwen3 235B A22B at Q4_K_M. Usable-memory figures are conservative: Windows shares about half of system RAM with the GPU by default (Linux can expose more), macOS lets Apple Silicon GPUs use about 75% of unified memory, and Linux unified-memory systems such as GB10 expose roughly 90%. Generation speed is bound by memory bandwidth, so compare the GB/s column before buying.

SystemTypeMemoryUsable for AIBandwidthGPUs
HP Z8 Fury G5 (4x RTX 6000 Ada) Workstation 192 GB 192 GB 960 GB/s 4

Frequently asked questions