AI GPU Benchmark Results

Every benchmark number in our database, with its source and tier. Filter by workload or source tier; sort by any column. Tier 1 = manufacturer/benchmark-DB, Tier 2 = independent reviews, Tier 3 = community or estimated.

97sourced results
26products covered
64GPUs in database
Honest coverage: our current benchmark set covers LLM inference (tokens/s) and Stable Diffusion image generation on flagship and mid-range cards. Gaming FPS and video-generation numbers are not yet sourced — they will appear here when we have verifiable sources. See the methodology and tier definitions.

Filter results

97 benchmark results

GPU Workload Model / config Result Source tier Source
NVIDIA B200 LLM inference (tokens/s) Llama-2-70B · FP8 · TensorRT-LLM (MLPerf v5.1) (MLPerf v5.1 server scenario, 8x B200 HGX host (throughput is per-host, not per-GPU)) 101,611 tok/s T1 MLPerf Inference Results
NVIDIA B200 LLM inference (tokens/s) Llama-2-70B · FP8 · TensorRT-LLM (MLPerf v5.1) (MLPerf v5.1 offline scenario, 8x B200 HGX host (throughput is per-host, not per-GPU)) 101,246 tok/s T1 MLPerf Inference Results
NVIDIA B200 LLM inference (tokens/s) Llama-3.1-405B · FP8 · TensorRT-LLM (MLPerf v5.1) (MLPerf v5.1 offline scenario, 8x B200 HGX host (throughput is per-host, not per-GPU)) 1,660 tok/s T1 MLPerf Inference Results
NVIDIA B200 LLM inference (tokens/s) Llama-3.1-405B · FP8 · TensorRT-LLM (MLPerf v5.1) (MLPerf v5.1 server scenario, 8x B200 HGX host (throughput is per-host, not per-GPU)) 1,280 tok/s T1 MLPerf Inference Results
H200 SXM LLM inference (tokens/s) Llama-3-70B · Q4 · llama.cpp (141GB HBM3e — editorial estimate pending article-level source verification) 140 tok/s EST T1 NVIDIA Official Specs
H100 SXM LLM inference (tokens/s) Llama-3-70B · Q4 · llama.cpp (80GB HBM3 — editorial estimate pending article-level source verification) 90 tok/s EST T1 MLPerf Inference Results
H100 SXM LLM inference (tokens/s) Llama-3-8B · Q4 · llama.cpp (80GB HBM3 — editorial estimate pending article-level source verification) 500 tok/s EST T1 NVIDIA Official Specs

How these numbers are sourced

Every row links to its original source. Tier 1 comes from manufacturers or standardized benchmark databases (MLPerf), Tier 2 from independent reviews (Tom's Hardware, TechPowerUp, Puget Systems), Tier 3 from community measurements or our own scaling estimates — always labeled EST when estimated. Read the full benchmark methodology.

Compare any two GPUs head-to-head in the comparison tool, or browse value rankings.