AI GPU Benchmark Results

Every benchmark number in our database, with its source and tier. Filter by workload or source tier; sort by any column. Tier 1 = manufacturer/benchmark-DB, Tier 2 = independent reviews, Tier 3 = community or estimated.

97sourced results
26products covered
64GPUs in database
Honest coverage: our current benchmark set covers LLM inference (tokens/s) and Stable Diffusion image generation on flagship and mid-range cards. Gaming FPS and video-generation numbers are not yet sourced — they will appear here when we have verifiable sources. See the methodology and tier definitions.

Filter results

GeForce RTX 5090 view all GPUs

97 benchmark results

GPU Workload Model / config Result Source tier Source
GeForce RTX 5090 LLM inference (tokens/s) Llama-2-13B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 134.502 tok/s) 6,591 score T2 StorageReview GPU Reviews
GeForce RTX 5090 LLM inference (tokens/s) Llama-3-70B · Q4 · llama.cpp (32GB VRAM, tight fit — editorial estimate pending article-level source verification) 42 tok/s EST T2 Tom's Hardware GPU Benchmarks 2025
GeForce RTX 5090 LLM inference (tokens/s) Llama-3-8B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 214.285 tok/s) 6,104 score T2 StorageReview GPU Reviews
GeForce RTX 5090 LLM inference (tokens/s) Llama-3.1-8B · FP16 · vLLM (Batch-serving throughput, ~18GB VRAM used; community + Spheron internal testing (approximate)) 3,500 tok/s EST T3 Spheron GPU Benchmark Blog
GeForce RTX 5090 LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) 215 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 5090 LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) 9,800 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 5090 LLM inference (tokens/s) Mistral-7B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 255.945 tok/s) 6,267 score T2 StorageReview GPU Reviews
GeForce RTX 5090 LLM inference (tokens/s) Phi-3.5-mini · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 314.435 tok/s) 5,749 score T2 StorageReview GPU Reviews
GeForce RTX 5090 LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 145.3 tok/s T3 Hardware Corner LLM GPU Rankings

How these numbers are sourced

Every row links to its original source. Tier 1 comes from manufacturers or standardized benchmark databases (MLPerf), Tier 2 from independent reviews (Tom's Hardware, TechPowerUp, Puget Systems), Tier 3 from community measurements or our own scaling estimates — always labeled EST when estimated. Read the full benchmark methodology.

Compare any two GPUs head-to-head in the comparison tool, or browse value rankings.