AI GPU Benchmark Results
Every benchmark number in our database, with its source and tier. Filter by workload or source tier; sort by any column. Tier 1 = manufacturer/benchmark-DB, Tier 2 = independent reviews, Tier 3 = community or estimated.
97sourced results
26products covered
64GPUs in database
Honest coverage: our current benchmark set covers LLM inference (tokens/s) and Stable Diffusion image generation on flagship and mid-range cards. Gaming FPS and video-generation numbers are not yet sourced — they will appear here when we have verifiable sources. See the methodology and tier definitions.
Filter results
97 benchmark results
| GPU | Workload | Model / config | Result | Source tier | Source |
|---|---|---|---|---|---|
| RTX A6000 | LLM inference (tokens/s) | Llama-3-8B · Q4_K_M · llama.cpp (CUDA, LLAMA_CUBLAS build) (Token generation: average speed generating 1024 tokens, batch 1, -ngl 10000 full offload, RunPod; model is Meta-Llama-3-8B (not 3.1); 2024-era build — cross-check: same repo's 4090=127.74 vs myaihardware b3500 125 (+2%); tested_at = repo last push containing Llama-3 results (exact run date unstated)) | 102.2 tok/s | T3 | XiongjieDai Multi-GPU llama.cpp Benchmarks |
| RTX A6000 | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 64.3 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| RTX 6000 Ada | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) | 110 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| RTX 6000 Ada | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) | 3,200 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| RTX 6000 Ada | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 98.7 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| Radeon RX 7900 XTX | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (ROCm) (Token generation: single user, batch 1, 4k context, mean of 3 runs) | 76 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| Radeon RX 7900 XTX | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (ROCm) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) | 2,100 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| NVIDIA RTX PRO 6000 Blackwell | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 140.6 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 5090 | LLM inference (tokens/s) | Llama-3.1-8B · FP16 · vLLM (Batch-serving throughput, ~18GB VRAM used; community + Spheron internal testing (approximate)) | 3,500 tok/s EST | T3 | Spheron GPU Benchmark Blog |
| GeForce RTX 5090 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) | 215 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 5090 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) | 9,800 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 5090 | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 145.3 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 5080 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp via Ollama (Token generation: single user, batch 1; source measured 5090=213 / 4090=127 on same rig, within 2% of myaihardware baseline) | 132 tok/s | T3 | LocalAI Master GPU Benchmarks |
| GeForce RTX 5080 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp via Ollama (Prompt processing: ~5,200 tok/s reported (approximate, batch 1)) | 5,200 tok/s | T3 | LocalAI Master GPU Benchmarks |
| GeForce RTX 5080 | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 94.1 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 5070 Ti | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 87.5 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 5070 | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 59.1 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 5060 Ti 16GB | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 51.4 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 4090 | LLM inference (tokens/s) | Llama-3.1-8B · FP16 · vLLM (Batch-serving throughput, ~18GB VRAM used; published llama.cpp/vLLM benchmarks (approximate)) | 2,550 tok/s EST | T3 | Spheron GPU Benchmark Blog |
| GeForce RTX 4090 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) | 125 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 4090 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) | 4,800 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 4090 | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 104.3 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 4080 SUPER | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) | 102 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 4080 SUPER | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) | 3,900 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 4080 SUPER | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 79.4 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 4080 | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 77.9 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 4070 Ti SUPER | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 72.2 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 4070 SUPER | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) | 92 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 4070 SUPER | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) | 2,900 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 4070 SUPER | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 56.2 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 4070 | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 52.1 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 4060 Ti 16GB | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 34.3 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 3090 Ti | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 93.6 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 3090 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) | 85 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 3090 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) | 2,400 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 3090 | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 87.5 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 3080 Ti | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 87.9 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 3080 10GB | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 74.2 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 3060 12GB | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) | 38 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 3060 12GB | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) | 950 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 3060 12GB | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 42 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| Arc B580 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp (Vulkan) (Token generation: single user, batch 1; source reports ~40–42 tok/s range (midpoint stored); Vulkan backend — SYCL/IPEX-LLM paths vary 15–38 tok/s on 14B models) | 41 tok/s | T3 | RunAIHome Local AI Benchmarks |
| Apple M3 Max | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (Metal) (Token generation: single user, batch 1, 4k context, mean of 3 runs) | 55 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| Apple M3 Max | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (Metal) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) | 1,200 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
How these numbers are sourced
Every row links to its original source. Tier 1 comes from manufacturers or standardized benchmark databases (MLPerf), Tier 2 from independent reviews (Tom's Hardware, TechPowerUp, Puget Systems), Tier 3 from community measurements or our own scaling estimates — always labeled EST when estimated. Read the full benchmark methodology.
Compare any two GPUs head-to-head in the comparison tool, or browse value rankings.