AI GPU Benchmark Results
Every benchmark number in our database, with its source and evidence type. Filter by workload or evidence type; sort by any column. Independent benchmarks come from third-party reviews and standardized databases (MLPerf); community results are labeled reproducible when the run config is published, estimated when approximated.
Filter results
100 benchmark results
| GPU | Workload | Model / config | Result | Evidence | Source |
|---|---|---|---|---|---|
| GeForce RTX 4090 | Image generation | FLUX.1-dev · BF16 · Diffusers (Requires memory-efficient attention (xFormers/SDPA) within 24GB; approximate) | 4 img/min EST | Estimated0.40 | Spheron GPU Benchmark Blog |
| GeForce RTX 4090 | LLM inference (tokens/s) | Llama-2-13B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 92.853 tok/s) | 5,013 score | Independent benchmark0.85 | StorageReview GPU Reviews |
| GeForce RTX 4090 | LLM inference (tokens/s) | Llama-3-70B · Q4 · llama.cpp (24GB VRAM, very tight, heavy swapping — editorial estimate pending article-level source verification) | 18 tok/s EST | Estimated0.40 | Tom's Hardware GPU Benchmarks 2025 |
| GeForce RTX 4090 | LLM inference (tokens/s) | Llama-3-8B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 150.039 tok/s) | 4,849 score | Independent benchmark0.85 | StorageReview GPU Reviews |
| GeForce RTX 4090 | LLM inference (tokens/s) | Llama-3.1-8B · FP16 · vLLM (Batch-serving throughput, ~18GB VRAM used; published llama.cpp/vLLM benchmarks (approximate)) | 2,550 tok/s EST | Estimated0.40 | Spheron GPU Benchmark Blog |
| GeForce RTX 4090 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) | 125 tok/s | Unverified community result0.55 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 4090 | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) | 4,800 tok/s | Unverified community result0.55 | MyAIHardware llama.cpp Benchmarks |
| GeForce RTX 4090 | LLM inference (tokens/s) | Mistral-7B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 183.266 tok/s) | 5,094 score | Independent benchmark0.85 | StorageReview GPU Reviews |
| GeForce RTX 4090 | LLM inference (tokens/s) | Phi-3.5-mini · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 244.343 tok/s) | 4,958 score | Independent benchmark0.85 | StorageReview GPU Reviews |
| GeForce RTX 4090 | LLM inference (tokens/s) | Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) | 104.3 tok/s | Unverified community result0.55 | Hardware Corner LLM GPU Rankings |
| GeForce RTX 4090 | Image generation | SDXL Turbo · FP16 · ComfyUI | 80 img/min | Independent benchmark0.85 | Tom's Hardware GPU Benchmarks 2025 |
| GeForce RTX 4090 | Image generation | Stable Diffusion 1.5 · FP16 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 1.188 s/image) | 5,260 score | Independent benchmark0.85 | StorageReview GPU Reviews |
| GeForce RTX 4090 | Image generation | Stable Diffusion 1.5 INT8 · INT8 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 0.503 s/image) | 62,160 score | Independent benchmark0.85 | StorageReview GPU Reviews |
| GeForce RTX 4090 | Image generation | Stable Diffusion XL · FP16 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 7.461 s/image) | 5,025 score | Independent benchmark0.85 | StorageReview GPU Reviews |
How these numbers are sourced
Every row links to its original source and carries an evidence type plus a confidence score. Independent benchmarks come from third-party reviews (Tom's Hardware, TechPowerUp, Puget Systems) and standardized databases (MLPerf); manufacturer benchmarks are vendor-published figures; community results are reproducible when the run config is published, unverified otherwise — always labeled EST when estimated. Read the full benchmark methodology.
Compare any two GPUs head-to-head in the comparison tool, or browse value rankings.