⌘K
← All AI Models

Can you run Deepseek R1 locally?

Deepseek family · · 684.5B parameters (37B activated per token) · Hugging Face model card

Deepseek R1 is a mixture-of-experts model with 37.0B active parameters per token and 684.5B total open-weight model with a 163,840-token context window, usable locally with roughly 512 GB of memory at Q4-class quantization (calculated; measured speeds vary by hardware).

Minimum: 512 GB+ memory (Q4_K_M, calculated)  ·  Recommended: 768 GB+ memory for full context (Q4_K_M, calculated)

Parameters (total)
684.5B
Activated per token
37B (MoE)
Context window
163,840 tokens
Architecture source
HF config.json (verified)

How much VRAM does Deepseek R1 need at each quantization?

Deepseek R1 needs 451.8 GB of VRAM at Q4_K_M. The table below lists weights-only size and total VRAM including overhead for each common quantization level with 4k and 32k token contexts.

QuantizationBits / weightWeights onlyTotal + KV @4k ctxTotal + KV @32k ctx
Q4_K_M 4.8 410.7 GB 451.8 GB 451.8 GB
Q5_K_M 5.7 487.7 GB 536.5 GB 536.5 GB
Q6_K 6.6 564.7 GB 621.2 GB 621.2 GB
Q8_0 8.5 727.3 GB 800.0 GB 800.0 GB
FP16 16 1,369.0 GB 1,505.9 GB 1,505.9 GB

Method: weights = parameters × bits-per-weight (Q4_K_M ≈ 4.8, Q5_K_M ≈ 5.7, Q6_K ≈ 6.6, Q8_0 ≈ 8.5, FP16 = 16), plus 10% loading overhead, plus KV cache from the verified architecture config (61 layers, 128 KV heads, head dim). Source: HF-verified 2026-09-12 via api (params=684.49B ctx=163840 license=mit); VRAM figures are documented calculations, not measurements

Which GPUs can run Deepseek R1 locally?

At Q4_K_M with a 4k context, Deepseek R1 needs 451.8 GB of VRAM. The lists below are computed live from our GPU database and grouped by how much headroom the card has. Disclosure: CompareAIHardware.com participates in the Amazon Associates program and earns from qualifying purchases through links on this page. Affiliate relationships do not influence our recommendations.

Needs 2+ GPUs to run Deepseek R1

One of these cards is too small on its own, but a pair (tensor or pipeline parallel, ~90% efficiency) covers the 451.8 GB requirement. See our multi-GPU guide for setup.

GPUVRAMBandwidthType
AMD Instinct MI325X 256 GB ×2 6000 GB/s Data Center GPU
AMD Instinct MI355X 288 GB ×2 8000 GB/s Data Center GPU
NVIDIA B300 (Blackwell Ultra) 288 GB ×2 8000 GB/s Data Center GPU

Frequently asked questions