Can you run DeepSeek V4 Flash Base locally?
DeepSeek V4 family · Apr 2026 (preview) · 284B parameters (13B activated per token) · Hugging Face model card
DeepSeek V4-Flash-Base is the base (pretrained) checkpoint of V4-Flash: 284B total / 13B activated, 1M context, official README table. Config matches Flash (deepseek-ai/DeepSeek-V4-Flash-Base@8855555deef230a27a21a8d6f294b7b7497759b6): 43 layers, hidden 4096, 64 heads, 1 KV head, head_dim=512, 256 routed experts × 6 + 1 shared. Vendor precision is FP8 Mixed (not the instruct FP4+FP8 mix). HF card license field empty; sibling instruct repo is MIT. Resident Q4_K_M ~187 GB (calculated from 284B). Measured tok/s: not published here.
Minimum: 1× NVIDIA B200 192GB (Q4_K_M ~187 GB, calculated from 284B resident) or 4× 80GB · Recommended: 2× NVIDIA B200 192GB or 8× H100 80GB
How much VRAM does DeepSeek V4 Flash Base need at each quantization?
DeepSeek V4 Flash Base needs 187.8 GB of VRAM at Q4_K_M. The table below lists weights-only size and total VRAM including overhead for each common quantization level with 4k and 32k token contexts.
| Quantization | Bits / weight | Weights only | Total + KV @4k ctx | Total + KV @32k ctx |
|---|---|---|---|---|
| Q4_K_M | 4.8 | 170.4 GB | 187.8 GB | 190.3 GB |
| Q5_K_M | 5.7 | 202.4 GB | 223.0 GB | 225.5 GB |
| Q6_K | 6.6 | 234.3 GB | 258.1 GB | 260.6 GB |
| Q8_0 | 8.5 | 301.8 GB | 332.3 GB | 334.8 GB |
| FP16 | 16 | 568.0 GB | 625.2 GB | 627.7 GB |
Same Flash architecture (config identical on published arch keys). VRAM follows 284B resident.
Method: weights = parameters × bits-per-weight (Q4_K_M ≈ 4.8, Q5_K_M ≈ 5.7, Q6_K ≈ 6.6, Q8_0 ≈ 8.5, FP16 = 16), plus 10% loading overhead, plus KV cache from the verified architecture config (43 layers, 1 KV heads, 512 head dim). Source: HF-verified 2026-09-21 via api (gated=false private=false safetensors.total=292021347282) + config.json repo=deepseek-ai/DeepSeek-V4-Flash-Base@8855555deef230a27a21a8d6f294b7b7497759b6 model_type=deepseek_v4 layers=43 hidden=4096 heads=64 kv_heads=1 head_dim=512 ctx=1048576 n_routed_experts=256 num_experts_per_tok=6 n_shared_experts=1 sliding_window=128 license=mit. Vendor README table: total=284B active=13B precision=fp8_mixed. VRAM GB are T08 documented calculations from resident total, not measurements. No tok/s invented.
Which GPUs can run DeepSeek V4 Flash Base locally?
At Q4_K_M with a 4k context, DeepSeek V4 Flash Base needs 187.8 GB of VRAM. The lists below are computed live from our GPU database and grouped by how much headroom the card has. Disclosure: CompareAIHardware.com participates in the Amazon Associates program and earns from qualifying purchases through links on this page. Affiliate relationships do not influence our recommendations.
Runs DeepSeek V4 Flash Base comfortably (25%+ VRAM headroom)
| GPU | VRAM | Bandwidth | Type | |
|---|---|---|---|---|
| AMD Instinct MI325X | 256 GB | 6000 GB/s | Data Center GPU | |
| AMD Instinct MI355X | 288 GB | 8000 GB/s | Data Center GPU | |
| NVIDIA B300 (Blackwell Ultra) | 288 GB | 8000 GB/s | Data Center GPU |
Minimum GPUs that fit DeepSeek V4 Flash Base at Q4_K_M
These GPUs hold the model but leave little headroom — keep contexts short.
| GPU | VRAM | Bandwidth | Type | |
|---|---|---|---|---|
| AMD Instinct MI300X | 192 GB | 5300 GB/s | Data Center GPU | |
| NVIDIA B200 | 192 GB | 8000 GB/s | Data Center GPU |
Needs 2+ GPUs to run DeepSeek V4 Flash Base
One of these cards is too small on its own, but a pair (tensor or pipeline parallel, ~90% efficiency) covers the 187.8 GB requirement. See our multi-GPU guide for setup.
| GPU | VRAM | Bandwidth | Type | |
|---|---|---|---|---|
| H200 SXM | 141 GB ×2 | 4800 GB/s | Data Center GPU |
Can a Mac run DeepSeek V4 Flash Base?
Yes — these Apple Silicon machines fit DeepSeek V4 Flash Base at Q4_K_M, since macOS lets the GPU use about 75% of unified memory. Generation speed is bound by memory bandwidth, so the GB/s column matters as much as capacity.
| Apple Silicon | Unified memory | Usable by GPU (~75%) | Bandwidth |
|---|---|---|---|
| M3 Ultra (Mac Studio, 256 GB) | 256 GB | 192 GB | 819 GB/s |
Which pre-built systems can run DeepSeek V4 Flash Base?
These mini PCs and workstations from our database fit DeepSeek V4 Flash Base at Q4_K_M. Usable-memory figures are conservative: Windows shares about half of system RAM with the GPU by default (Linux can expose more), macOS lets Apple Silicon GPUs use about 75% of unified memory, and Linux unified-memory systems such as GB10 expose roughly 90%. Generation speed is bound by memory bandwidth, so compare the GB/s column before buying.
| System | Type | Memory | Usable for AI | Bandwidth | GPUs |
|---|---|---|---|---|---|
| HP Z8 Fury G5 (4x RTX 6000 Ada) | Workstation | 192 GB | 192 GB | 960 GB/s | 4 |