AI GPU Benchmark Results
Every benchmark number in our database, with its source and evidence type. Filter by workload or evidence type; sort by any column. Independent benchmarks come from third-party reviews and standardized databases (MLPerf); community results are labeled reproducible when the run config is published, estimated when approximated.
Filter results
100 benchmark results
| GPU | Workload | Model / config | Result | Evidence | Source |
|---|---|---|---|---|---|
| RTX A6000 | LLM inference (tokens/s) | Llama-3-70B · Q4 · llama.cpp (48GB VRAM — editorial estimate pending article-level source verification) | 22 tok/s EST | Estimated0.40 | Puget Systems Hardware Testing |
| RTX A6000 | LLM inference (tokens/s) | Llama-3-8B · Q4_K_M · llama.cpp (CUDA, LLAMA_CUBLAS build) (Token generation: average speed generating 1024 tokens, batch 1, -ngl 10000 full offload, RunPod; model is Meta-Llama-3-8B (not 3.1); 2024-era build — cross-check: same repo's 4090=127.74 vs myaihardware b3500 125 (+2%); tested_at = repo last push containing Llama-3 results (exact run date unstated)) | 102.2 tok/s | Reproducible community benchmark0.70 | XiongjieDai Multi-GPU llama.cpp Benchmarks |
| RTX A6000 | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 64.3 tok/s | Unverified community result0.55 | Hardware Corner LLM GPU Rankings |
How these numbers are sourced
Every row links to its original source and carries an evidence type plus a confidence score. Independent benchmarks come from third-party reviews (Tom's Hardware, TechPowerUp, Puget Systems) and standardized databases (MLPerf); manufacturer benchmarks are vendor-published figures; community results are reproducible when the run config is published, unverified otherwise — always labeled EST when estimated. Read the full benchmark methodology.
Compare any two GPUs head-to-head in the comparison tool, or browse value rankings.