← All GPUs
NVIDIA B200
NVIDIA · Data Center GPU · 2025
Performance Benchmarks
LLM Inference all sourced rows →
| Model | Metric | Value | Source |
|---|---|---|---|
|
Llama-2-70B
FP8 MLPerf v5.1 server scenario, 8x B200 HGX host (throughput is per-host, not per-GPU) |
system tokens per second server |
101,611
tok/s
TensorRT-LLM (MLPerf v5.1) · 2025-09-10
|
Verified MLPerf Inference Results ↗ |
|
Llama-2-70B
FP8 MLPerf v5.1 offline scenario, 8x B200 HGX host (throughput is per-host, not per-GPU) |
system tokens per second offline |
101,246
tok/s
TensorRT-LLM (MLPerf v5.1) · 2025-09-10
|
Verified MLPerf Inference Results ↗ |
|
Llama-3.1-405B
FP8 MLPerf v5.1 offline scenario, 8x B200 HGX host (throughput is per-host, not per-GPU) |
system tokens per second offline |
1,660
tok/s
TensorRT-LLM (MLPerf v5.1) · 2025-09-10
|
Verified MLPerf Inference Results ↗ |
|
Llama-3.1-405B
FP8 MLPerf v5.1 server scenario, 8x B200 HGX host (throughput is per-host, not per-GPU) |
system tokens per second server |
1,280
tok/s
TensorRT-LLM (MLPerf v5.1) · 2025-09-10
|
Verified MLPerf Inference Results ↗ |
Full Specifications
Vram gb
192 GB
Vram type
HBM3e
Memory bandwidth gbps
8000 GB/s
Tdp watts
1000 W
Architecture
Blackwell
Interconnect
NVLink 5 (1.8 TB/s)
Related GPUs
Guides & Tools
CompareAIHardware.com participates in the Amazon Associates program. As an Amazon Associate, we earn from qualifying purchases.