Can you run Llama 3.1 70B locally?
Llama 3 family · July 2024 · 70.60B parameters · Hugging Face model card
You need a GPU with 48 GB of VRAM to run Llama 3.1 70B locally at Q4_K_M with a 4k-token context, which requires about 47.9 GB of VRAM. An 80 GB data center GPU such as the NVIDIA H100 80GB or A100 80GB runs the model comfortably, and a 32k-token context raises the requirement to about 57.3 GB of VRAM.
Minimum: 48 GB VRAM GPU (RTX A6000, RTX 6000 Ada, L40S) at Q4_K_M, 4k context · Recommended: 80 GB VRAM GPU (H100 80GB, A100 80GB) at Q5_K_M–Q6_K for full context
How much VRAM does Llama 3.1 70B need at each quantization?
Llama 3.1 70B needs 47.9 GB of VRAM at Q4_K_M. The table below lists weights-only size and total VRAM including overhead for each common quantization level with 4k and 32k token contexts.
| Quantization | Bits / weight | Weights only | Total + KV @4k ctx | Total + KV @32k ctx |
|---|---|---|---|---|
| Q4_K_M | 4.8 | 42.4 GB | 47.9 GB | 57.3 GB |
| Q5_K_M | 5.7 | 50.3 GB | 56.7 GB | 66.1 GB |
| Q6_K | 6.6 | 58.3 GB | 65.4 GB | 74.8 GB |
| Q8_0 | 8.5 | 75.0 GB | 83.9 GB | 93.3 GB |
| FP16 | 16 | 141.2 GB | 156.7 GB | 166.1 GB |
Method: weights = parameters × bits-per-weight (Q4_K_M ≈ 4.8, Q5_K_M ≈ 5.7, Q6_K ≈ 6.6, Q8_0 ≈ 8.5, FP16 = 16), plus 10% loading overhead, plus KV cache from the verified architecture config (80 layers, 8 KV heads, 128 head dim). Source: Params 70.6B, 80 layers, 8 KV heads, 128 head dim, 131072 context — HF config.json (NousResearch/Meta-Llama-3.1-70B-Instruct).
Which GPUs can run Llama 3.1 70B locally?
At Q4_K_M with a 4k context, Llama 3.1 70B needs 47.9 GB of VRAM. The lists below are computed live from our GPU database and grouped by how much headroom the card has. Disclosure: CompareAIHardware.com participates in the Amazon Associates program and earns from qualifying purchases through links on this page. Affiliate relationships do not influence our recommendations.
Runs Llama 3.1 70B comfortably (25%+ VRAM headroom)
| GPU | VRAM | Bandwidth | Type | |
|---|---|---|---|---|
| H100 SXM | 80 GB | 3350 GB/s | Data Center GPU | |
| NVIDIA A100 80GB SXM | 80 GB | 2039 GB/s | Data Center GPU | |
| NVIDIA H100 PCIe 80GB | 80 GB | 2039 GB/s | Data Center GPU | |
| NVIDIA RTX PRO 6000 Blackwell | 96 GB | 1792 GB/s | Pro GPU | Check price |
| H200 SXM | 141 GB | 4800 GB/s | Data Center GPU | |
| AMD Instinct MI300X | 192 GB | 5300 GB/s | Data Center GPU | |
| NVIDIA B200 | 192 GB | 8000 GB/s | Data Center GPU | |
| AMD Instinct MI325X | 256 GB | 6000 GB/s | Data Center GPU | |
| AMD Instinct MI355X | 288 GB | 8000 GB/s | Data Center GPU | |
| NVIDIA B300 (Blackwell Ultra) | 288 GB | 8000 GB/s | Data Center GPU |
Minimum GPUs that fit Llama 3.1 70B at Q4_K_M
These GPUs hold the model but leave little headroom — keep contexts short.
| GPU | VRAM | Bandwidth | Type | |
|---|---|---|---|---|
| AMD Radeon Pro W7900 | 48 GB | 864 GB/s | Pro GPU | Check price |
| NVIDIA RTX PRO 5000 Blackwell | 48 GB | 1344 GB/s | Pro GPU | |
| RTX 6000 Ada | 48 GB | 960 GB/s | Pro GPU | Check price |
| RTX A6000 | 48 GB | 768 GB/s | Pro GPU | Check price |
Needs 2+ GPUs to run Llama 3.1 70B
One of these cards is too small on its own, but a pair (tensor or pipeline parallel, ~90% efficiency) covers the 47.9 GB requirement. See our multi-GPU guide for setup.
| GPU | VRAM | Bandwidth | Type | |
|---|---|---|---|---|
| AMD Radeon Pro W7800 | 32 GB ×2 | 576 GB/s | Pro GPU | Check price |
| GeForce RTX 5090 | 32 GB ×2 | 1792 GB/s | Consumer GPU | Check price |
| NVIDIA RTX 5000 Ada | 32 GB ×2 | 576 GB/s | Pro GPU | Check price |
| NVIDIA RTX PRO 4500 Blackwell | 32 GB ×2 | 896 GB/s | Pro GPU | Check price |
| NVIDIA A100 40GB SXM EOL | 40 GB ×2 | 1555 GB/s | Data Center GPU |
Can a Mac run Llama 3.1 70B?
Yes — these Apple Silicon machines fit Llama 3.1 70B at Q4_K_M, since macOS lets the GPU use about 75% of unified memory. Generation speed is bound by memory bandwidth, so the GB/s column matters as much as capacity.
| Apple Silicon | Unified memory | Usable by GPU (~75%) | Bandwidth |
|---|---|---|---|
| Apple M4 Pro | 64 GB | 48 GB | 273 GB/s |
| Apple M3 Max | 128 GB | 96 GB | 400 GB/s |
| M4 Max (MacBook Pro) | 128 GB | 96 GB | 546 GB/s |
| M3 Ultra (Mac Studio) | 512 GB | 384 GB | 819 GB/s |
Can Llama 3.1 70B run on mini PCs, Jetson, or NPU devices?
These edge and NPU devices from our database have enough memory for Llama 3.1 70B at Q4_K_M. Their memory bandwidth is far below discrete GPUs, so expect a fraction of desktop generation speed.
| Device | Memory | Bandwidth | Type |
|---|---|---|---|
| Jetson AGX Orin 64GB | 64 GB | 204 GB/s | Edge Compute Device |
| Snapdragon X Elite (X1E-84-100) | 64 GB | 135 GB/s | NPU Chip |
| AMD Ryzen AI 9 HX 370 (Strix Point) | 96 GB | 120 GB/s | NPU Chip |
Which pre-built systems can run Llama 3.1 70B?
These mini PCs and workstations from our database fit Llama 3.1 70B at Q4_K_M. Usable-memory figures are conservative: Windows shares about half of system RAM with the GPU by default (Linux can expose more), macOS lets Apple Silicon GPUs use about 75% of unified memory, and Linux unified-memory systems such as GB10 expose roughly 90%. Generation speed is bound by memory bandwidth, so compare the GB/s column before buying.
| System | Type | Memory | Usable for AI | Bandwidth | GPUs |
|---|---|---|---|---|---|
| ASRock NUC BOX-255H (Core Ultra 7 255H) | Mini PC | 96 GB | 48 GB | 102 GB/s | — |
| Apple Mac Studio M2 Ultra (64GB) | Workstation | 64 GB | 48 GB | 800 GB/s | 1 |
| Custom Dual RTX 4090 Training Workstation | Workstation | 48 GB | 48 GB | 1008 GB/s | 2 |
| Dell Precision 7960 Tower (1x RTX 6000 Ada) | Workstation | 48 GB | 48 GB | 960 GB/s | 1 |
| HP Z8 Fury G5 (1x RTX 6000 Ada) | Workstation | 48 GB | 48 GB | 960 GB/s | 1 |
| AMD Ryzen AI Halo Developer Platform (Max+ 395) | Mini PC | 128 GB | 64 GB | 256 GB/s | — |
| ArsenalPC MES2X Dual RTX 5090 AI Workstation | Workstation | 64 GB | 64 GB | — | 2 |
| GMKtec EVO-X2 (Ryzen AI Max+ 395) | Mini PC | 128 GB | 64 GB | 256 GB/s | — |
| Mac mini M4 Max (128GB) | Mini PC | 128 GB | 64 GB | 546 GB/s | — |
| MinisForum MS-S1 MAX (Ryzen AI Max+ 395) | Mini PC | 128 GB | 64 GB | 256 GB/s | — |
| BOXX APEXX 8R (1x RTX PRO 6000 Blackwell) | Workstation | 96 GB | 96 GB | 1792 GB/s | 1 |
| GEEKOM A9 Mega AI Workstation | Workstation | 128 GB | 96 GB | — | 1 |
| NOVATECH RTX PRO 6000 AI Workstation | Workstation | 96 GB | 96 GB | — | 1 |
| System76 Thelio Major (1x RTX PRO 6000 Blackwell) | Workstation | 96 GB | 96 GB | 1792 GB/s | 1 |
| System76 Thelio Major (2x RTX 6000 Ada) | Workstation | 96 GB | 96 GB | 960 GB/s | 2 |
| Lenovo ThinkStation PGX | Workstation | 128 GB | 115.2 GB | 273 GB/s | 1 |
| NVIDIA DGX Spark | Workstation | 128 GB | 115.2 GB | 273 GB/s | 1 |
| BIZON G3000 G2 (4x RTX 5090) | Workstation | 128 GB | 128 GB | 1792 GB/s | 4 |
| Apple Mac Studio M2 Ultra (192GB) | Workstation | 192 GB | 144 GB | 800 GB/s | 1 |
| HP Z8 Fury G5 (4x RTX 6000 Ada) | Workstation | 192 GB | 192 GB | 960 GB/s | 4 |
Frequently asked questions
Can an RTX 4090 run Llama 3.1 70B?
No. A single RTX 4090 with 24 GB of VRAM cannot hold the 42.4 GB Q4_K_M weights of Llama 3.1 70B. Two RTX 4090s reach 48 GB total but still fall short of the 47.9 GB requirement once overhead and KV cache are counted, so partial CPU offloading is required.
How much VRAM does Llama 3.1 70B need at Q4?
Llama 3.1 70B needs about 47.9 GB of VRAM at Q4_K_M with a 4k-token context: 42.36 GB of weights plus 10% loading overhead plus a 1.34 GB KV cache.
Can a Mac run Llama 3.1 70B?
Yes. A Mac Studio with M3 Ultra and 192 GB or more of unified memory runs Llama 3.1 70B at Q4_K_M, because macOS exposes about 75% of unified memory to the GPU. Macs with less than 64 GB of unified memory cannot hold the quantized weights.
Related guides
- Open Llama 3.1 70B in the Model-to-GPU finder — pick quantization and context, see ranked GPU matches.
- Best GPUs for LLMs
- Multi-GPU Guide
- VRAM Calculator