AI GPU Benchmark Results
Every benchmark number in our database, with its source and tier. Filter by workload or source tier; sort by any column. Tier 1 = manufacturer/benchmark-DB, Tier 2 = independent reviews, Tier 3 = community or estimated.
97sourced results
26products covered
64GPUs in database
Honest coverage: our current benchmark set covers LLM inference (tokens/s) and Stable Diffusion image generation on flagship and mid-range cards. Gaming FPS and video-generation numbers are not yet sourced — they will appear here when we have verifiable sources. See the methodology and tier definitions.
Filter results
97 benchmark results
| GPU | Workload | Model / config | Result | Source tier | Source |
|---|---|---|---|---|---|
| GeForce RTX 5080 | LLM inference (tokens/s) | Llama-2-13B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 83.653 tok/s) | 4,790 score | T2 | StorageReview GPU Reviews |
| GeForce RTX 5080 | LLM inference (tokens/s) | Llama-3-8B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 136.177 tok/s) | 4,424 score | T2 | StorageReview GPU Reviews |
| GeForce RTX 5080 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp via Ollama (Token generation: single user, batch 1; source measured 5090=213 / 4090=127 on same rig, within 2% of myaihardware baseline) | 132 tok/s | T3 | LocalAI Master GPU Benchmarks |
| GeForce RTX 5080 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp via Ollama (Prompt processing: ~5,200 tok/s reported (approximate, batch 1)) | 5,200 tok/s | T3 | LocalAI Master GPU Benchmarks |
| GeForce RTX 5080 | LLM inference (tokens/s) | Mistral-7B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 163.598 tok/s) | 4,635 score | T2 | StorageReview GPU Reviews |
| GeForce RTX 5080 | LLM inference (tokens/s) | Phi-3.5-mini · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 209.459 tok/s) | 4,400 score | T2 | StorageReview GPU Reviews |
| GeForce RTX 5080 | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 94.1 tok/s | T3 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 5080 | Image generation | SDXL Turbo · FP16 · ComfyUI | 65 img/min | T2 | TechPowerUp GPU Reviews |
| GeForce RTX 5080 | Image generation | Stable Diffusion 1.5 · FP16 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 1.344 s/image) | 4,650 score | T2 | StorageReview GPU Reviews |
| GeForce RTX 5080 | Image generation | Stable Diffusion 1.5 INT8 · INT8 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 0.561 s/image) | 55,683 score | T2 | StorageReview GPU Reviews |
| GeForce RTX 5080 | Image generation | Stable Diffusion XL · FP16 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 8.808 s/image) | 4,257 score | T2 | StorageReview GPU Reviews |
How these numbers are sourced
Every row links to its original source. Tier 1 comes from manufacturers or standardized benchmark databases (MLPerf), Tier 2 from independent reviews (Tom's Hardware, TechPowerUp, Puget Systems), Tier 3 from community measurements or our own scaling estimates — always labeled EST when estimated. Read the full benchmark methodology.
Compare any two GPUs head-to-head in the comparison tool, or browse value rankings.