← All GPUs

NVIDIA B200

NVIDIA · Data Center GPU · 2025

192 GB VRAM
8000 GB/s
1000W TDP

Performance Benchmarks

LLM Inference all sourced rows →

ModelMetricValueSource
Llama-2-70B FP8
MLPerf v5.1 server scenario, 8x B200 HGX host (throughput is per-host, not per-GPU)
system tokens per second server 101,611 tok/s
TensorRT-LLM (MLPerf v5.1) · 2025-09-10
Verified MLPerf Inference Results ↗
Llama-2-70B FP8
MLPerf v5.1 offline scenario, 8x B200 HGX host (throughput is per-host, not per-GPU)
system tokens per second offline 101,246 tok/s
TensorRT-LLM (MLPerf v5.1) · 2025-09-10
Verified MLPerf Inference Results ↗
Llama-3.1-405B FP8
MLPerf v5.1 offline scenario, 8x B200 HGX host (throughput is per-host, not per-GPU)
system tokens per second offline 1,660 tok/s
TensorRT-LLM (MLPerf v5.1) · 2025-09-10
Verified MLPerf Inference Results ↗
Llama-3.1-405B FP8
MLPerf v5.1 server scenario, 8x B200 HGX host (throughput is per-host, not per-GPU)
system tokens per second server 1,280 tok/s
TensorRT-LLM (MLPerf v5.1) · 2025-09-10
Verified MLPerf Inference Results ↗

Full Specifications

Vram gb
192 GB
Vram type
HBM3e
Memory bandwidth gbps
8000 GB/s
Tdp watts
1000 W
Architecture
Blackwell
Interconnect
NVLink 5 (1.8 TB/s)

Related GPUs

Guides & Tools

CompareAIHardware.com participates in the Amazon Associates program. As an Amazon Associate, we earn from qualifying purchases.