AI GPU Benchmark Results

Every benchmark number in our database, with its source and tier. Filter by workload or source tier; sort by any column. Tier 1 = manufacturer/benchmark-DB, Tier 2 = independent reviews, Tier 3 = community or estimated.

97sourced results
26products covered
64GPUs in database
Honest coverage: our current benchmark set covers LLM inference (tokens/s) and Stable Diffusion image generation on flagship and mid-range cards. Gaming FPS and video-generation numbers are not yet sourced — they will appear here when we have verifiable sources. See the methodology and tier definitions.

Filter results

97 benchmark results

GPU Workload Model / config Result Source tier Source
Apple M3 Max LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (Metal) (Token generation: single user, batch 1, 4k context, mean of 3 runs) 55 tok/s T3 MyAIHardware llama.cpp Benchmarks
Apple M3 Max LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (Metal) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) 1,200 tok/s T3 MyAIHardware llama.cpp Benchmarks
Arc B580 LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp (Vulkan) (Token generation: single user, batch 1; source reports ~40–42 tok/s range (midpoint stored); Vulkan backend — SYCL/IPEX-LLM paths vary 15–38 tok/s on 14B models) 41 tok/s T3 RunAIHome Local AI Benchmarks
Arc B580 Image generation SDXL Turbo · FP16 · ComfyUI 18 img/min EST T2 TechPowerUp GPU Reviews
GeForce RTX 3060 12GB LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) 38 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 3060 12GB LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) 950 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 3060 12GB LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 42 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 3080 10GB LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 74.2 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 3080 Ti LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 87.9 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 3090 LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) 85 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 3090 LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) 2,400 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 3090 LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 87.5 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 3090 Image generation SDXL Turbo · FP16 · ComfyUI 45 img/min T2 Puget Systems Hardware Testing
GeForce RTX 3090 Ti LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 93.6 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 4060 Ti 16GB LLM inference (tokens/s) Llama-3-8B · Q4 · llama.cpp (16GB VRAM — EST pending article-level Llama-3.1-8B verification; verified Qwen3-8B @16K row exists (34.31 tok/s) — editorial estimate pending article-level source verification) 95 tok/s EST T2 TechPowerUp GPU Reviews
GeForce RTX 4060 Ti 16GB LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 34.3 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 4060 Ti 16GB Image generation SDXL Turbo · FP16 · ComfyUI 28 img/min T2 TechPowerUp GPU Reviews
GeForce RTX 4070 LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 52.1 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 4070 SUPER LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) 92 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 4070 SUPER LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) 2,900 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 4070 SUPER LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 56.2 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 4070 Ti SUPER LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 72.2 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 4080 LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 77.9 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 4080 SUPER LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) 102 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 4080 SUPER LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) 3,900 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 4080 SUPER LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 79.4 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 4080 SUPER Image generation SDXL Turbo · FP16 · ComfyUI 58 img/min T2 TechPowerUp GPU Reviews
GeForce RTX 4090 Image generation FLUX.1-dev · BF16 · Diffusers (Requires memory-efficient attention (xFormers/SDPA) within 24GB; approximate) 4 img/min EST T3 Spheron GPU Benchmark Blog
GeForce RTX 4090 LLM inference (tokens/s) Llama-2-13B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 92.853 tok/s) 5,013 score T2 StorageReview GPU Reviews
GeForce RTX 4090 LLM inference (tokens/s) Llama-3-70B · Q4 · llama.cpp (24GB VRAM, very tight, heavy swapping — editorial estimate pending article-level source verification) 18 tok/s EST T2 Tom's Hardware GPU Benchmarks 2025
GeForce RTX 4090 LLM inference (tokens/s) Llama-3-8B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 150.039 tok/s) 4,849 score T2 StorageReview GPU Reviews
GeForce RTX 4090 LLM inference (tokens/s) Llama-3.1-8B · FP16 · vLLM (Batch-serving throughput, ~18GB VRAM used; published llama.cpp/vLLM benchmarks (approximate)) 2,550 tok/s EST T3 Spheron GPU Benchmark Blog
GeForce RTX 4090 LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) 125 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 4090 LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) 4,800 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 4090 LLM inference (tokens/s) Mistral-7B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 183.266 tok/s) 5,094 score T2 StorageReview GPU Reviews
GeForce RTX 4090 LLM inference (tokens/s) Phi-3.5-mini · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 244.343 tok/s) 4,958 score T2 StorageReview GPU Reviews
GeForce RTX 4090 LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 104.3 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 4090 Image generation SDXL Turbo · FP16 · ComfyUI 80 img/min T2 Tom's Hardware GPU Benchmarks 2025
GeForce RTX 4090 Image generation Stable Diffusion 1.5 · FP16 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 1.188 s/image) 5,260 score T2 StorageReview GPU Reviews
GeForce RTX 4090 Image generation Stable Diffusion 1.5 INT8 · INT8 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 0.503 s/image) 62,160 score T2 StorageReview GPU Reviews
GeForce RTX 4090 Image generation Stable Diffusion XL · FP16 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 7.461 s/image) 5,025 score T2 StorageReview GPU Reviews
GeForce RTX 5060 Ti 16GB LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 51.4 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 5070 LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 59.1 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 5070 Ti LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 87.5 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 5080 LLM inference (tokens/s) Llama-2-13B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 83.653 tok/s) 4,790 score T2 StorageReview GPU Reviews
GeForce RTX 5080 LLM inference (tokens/s) Llama-3-8B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 136.177 tok/s) 4,424 score T2 StorageReview GPU Reviews
GeForce RTX 5080 LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp via Ollama (Token generation: single user, batch 1; source measured 5090=213 / 4090=127 on same rig, within 2% of myaihardware baseline) 132 tok/s T3 LocalAI Master GPU Benchmarks
GeForce RTX 5080 LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp via Ollama (Prompt processing: ~5,200 tok/s reported (approximate, batch 1)) 5,200 tok/s T3 LocalAI Master GPU Benchmarks
GeForce RTX 5080 LLM inference (tokens/s) Mistral-7B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 163.598 tok/s) 4,635 score T2 StorageReview GPU Reviews
GeForce RTX 5080 LLM inference (tokens/s) Phi-3.5-mini · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 209.459 tok/s) 4,400 score T2 StorageReview GPU Reviews
GeForce RTX 5080 LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 94.1 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 5080 Image generation SDXL Turbo · FP16 · ComfyUI 65 img/min T2 TechPowerUp GPU Reviews
GeForce RTX 5080 Image generation Stable Diffusion 1.5 · FP16 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 1.344 s/image) 4,650 score T2 StorageReview GPU Reviews
GeForce RTX 5080 Image generation Stable Diffusion 1.5 INT8 · INT8 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 0.561 s/image) 55,683 score T2 StorageReview GPU Reviews
GeForce RTX 5080 Image generation Stable Diffusion XL · FP16 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 8.808 s/image) 4,257 score T2 StorageReview GPU Reviews
GeForce RTX 5090 Image generation FLUX.1-dev · BF16 · Diffusers (~26GB VRAM used; community + Spheron internal testing (approximate)) 5.5 img/min EST T3 Spheron GPU Benchmark Blog
GeForce RTX 5090 LLM inference (tokens/s) Llama-2-13B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 134.502 tok/s) 6,591 score T2 StorageReview GPU Reviews
GeForce RTX 5090 LLM inference (tokens/s) Llama-3-70B · Q4 · llama.cpp (32GB VRAM, tight fit — editorial estimate pending article-level source verification) 42 tok/s EST T2 Tom's Hardware GPU Benchmarks 2025
GeForce RTX 5090 LLM inference (tokens/s) Llama-3-8B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 214.285 tok/s) 6,104 score T2 StorageReview GPU Reviews
GeForce RTX 5090 LLM inference (tokens/s) Llama-3.1-8B · FP16 · vLLM (Batch-serving throughput, ~18GB VRAM used; community + Spheron internal testing (approximate)) 3,500 tok/s EST T3 Spheron GPU Benchmark Blog
GeForce RTX 5090 LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) 215 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 5090 LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) 9,800 tok/s T3 MyAIHardware llama.cpp Benchmarks
GeForce RTX 5090 LLM inference (tokens/s) Mistral-7B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 255.945 tok/s) 6,267 score T2 StorageReview GPU Reviews
GeForce RTX 5090 LLM inference (tokens/s) Phi-3.5-mini · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 314.435 tok/s) 5,749 score T2 StorageReview GPU Reviews
GeForce RTX 5090 LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 145.3 tok/s T3 Hardware Corner LLM GPU Rankings
GeForce RTX 5090 Image generation SDXL Turbo · FP16 · ComfyUI 120 img/min T2 Tom's Hardware GPU Benchmarks 2025
GeForce RTX 5090 Image generation Stable Diffusion 1.5 · FP16 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 0.763 s/image) 8,193 score T2 StorageReview GPU Reviews
GeForce RTX 5090 Image generation Stable Diffusion 1.5 INT8 · INT8 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 0.394 s/image) 79,272 score T2 StorageReview GPU Reviews
GeForce RTX 5090 Image generation Stable Diffusion XL · FP16 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 5.223 s/image) 7,179 score T2 StorageReview GPU Reviews
H100 SXM LLM inference (tokens/s) Llama-3-70B · Q4 · llama.cpp (80GB HBM3 — editorial estimate pending article-level source verification) 90 tok/s EST T1 MLPerf Inference Results
H100 SXM LLM inference (tokens/s) Llama-3-8B · Q4 · llama.cpp (80GB HBM3 — editorial estimate pending article-level source verification) 500 tok/s EST T1 NVIDIA Official Specs
H100 SXM Image generation SDXL Turbo · FP16 · ComfyUI 200 img/min T1 NVIDIA Official Specs
H200 SXM LLM inference (tokens/s) Llama-3-70B · Q4 · llama.cpp (141GB HBM3e — editorial estimate pending article-level source verification) 140 tok/s EST T1 NVIDIA Official Specs
NVIDIA B200 LLM inference (tokens/s) Llama-2-70B · FP8 · TensorRT-LLM (MLPerf v5.1) (MLPerf v5.1 server scenario, 8x B200 HGX host (throughput is per-host, not per-GPU)) 101,611 tok/s T1 MLPerf Inference Results
NVIDIA B200 LLM inference (tokens/s) Llama-2-70B · FP8 · TensorRT-LLM (MLPerf v5.1) (MLPerf v5.1 offline scenario, 8x B200 HGX host (throughput is per-host, not per-GPU)) 101,246 tok/s T1 MLPerf Inference Results
NVIDIA B200 LLM inference (tokens/s) Llama-3.1-405B · FP8 · TensorRT-LLM (MLPerf v5.1) (MLPerf v5.1 offline scenario, 8x B200 HGX host (throughput is per-host, not per-GPU)) 1,660 tok/s T1 MLPerf Inference Results
NVIDIA B200 LLM inference (tokens/s) Llama-3.1-405B · FP8 · TensorRT-LLM (MLPerf v5.1) (MLPerf v5.1 server scenario, 8x B200 HGX host (throughput is per-host, not per-GPU)) 1,280 tok/s T1 MLPerf Inference Results
NVIDIA RTX PRO 6000 Blackwell LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 140.6 tok/s T3 Hardware Corner LLM GPU Rankings
Radeon RX 7900 XTX LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (ROCm) (Token generation: single user, batch 1, 4k context, mean of 3 runs) 76 tok/s T3 MyAIHardware llama.cpp Benchmarks
Radeon RX 7900 XTX LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (ROCm) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) 2,100 tok/s T3 MyAIHardware llama.cpp Benchmarks
Radeon RX 7900 XTX Image generation SDXL Turbo · FP16 · ComfyUI (ROCm) 40 img/min T2 Tom's Hardware GPU Benchmarks 2025
RTX 6000 Ada LLM inference (tokens/s) Llama-2-13B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 78.532 tok/s) 3,957 score T2 StorageReview GPU Reviews
RTX 6000 Ada LLM inference (tokens/s) Llama-3-70B · Q4 · llama.cpp (48GB VRAM — editorial estimate pending article-level source verification) 35 tok/s EST T2 Puget Systems Hardware Testing
RTX 6000 Ada LLM inference (tokens/s) Llama-3-8B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 138.62 tok/s) 4,026 score T2 StorageReview GPU Reviews
RTX 6000 Ada LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) 110 tok/s T3 MyAIHardware llama.cpp Benchmarks
RTX 6000 Ada LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) 3,200 tok/s T3 MyAIHardware llama.cpp Benchmarks
RTX 6000 Ada LLM inference (tokens/s) Mistral-7B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 166.633 tok/s) 4,255 score T2 StorageReview GPU Reviews
RTX 6000 Ada LLM inference (tokens/s) Phi-3.5-mini · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 228.359 tok/s) 4,508 score T2 StorageReview GPU Reviews
RTX 6000 Ada LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 98.7 tok/s T3 Hardware Corner LLM GPU Rankings
RTX 6000 Ada Image generation SDXL Turbo · FP16 · ComfyUI 75 img/min T2 Puget Systems Hardware Testing
RTX 6000 Ada Image generation Stable Diffusion 1.5 · FP16 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 1.477 s/image) 4,230 score T2 StorageReview GPU Reviews
RTX 6000 Ada Image generation Stable Diffusion 1.5 INT8 · INT8 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 0.559 s/image) 55,901 score T2 StorageReview GPU Reviews
RTX 6000 Ada Image generation Stable Diffusion XL · FP16 · UL Procyon AI Image Generation (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 12.323 s/image) 3,043 score T2 StorageReview GPU Reviews
RTX A6000 LLM inference (tokens/s) Llama-3-70B · Q4 · llama.cpp (48GB VRAM — editorial estimate pending article-level source verification) 22 tok/s EST T2 Puget Systems Hardware Testing
RTX A6000 LLM inference (tokens/s) Llama-3-8B · Q4_K_M · llama.cpp (CUDA, LLAMA_CUBLAS build) (Token generation: average speed generating 1024 tokens, batch 1, -ngl 10000 full offload, RunPod; model is Meta-Llama-3-8B (not 3.1); 2024-era build — cross-check: same repo's 4090=127.74 vs myaihardware b3500 125 (+2%); tested_at = repo last push containing Llama-3 results (exact run date unstated)) 102.2 tok/s T3 XiongjieDai Multi-GPU llama.cpp Benchmarks
RTX A6000 LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 64.3 tok/s T3 Hardware Corner LLM GPU Rankings
RTX A6000 Image generation SDXL Turbo · FP16 · ComfyUI 35 img/min T2 Puget Systems Hardware Testing

How these numbers are sourced

Every row links to its original source. Tier 1 comes from manufacturers or standardized benchmark databases (MLPerf), Tier 2 from independent reviews (Tom's Hardware, TechPowerUp, Puget Systems), Tier 3 from community measurements or our own scaling estimates — always labeled EST when estimated. Read the full benchmark methodology.

Compare any two GPUs head-to-head in the comparison tool, or browse value rankings.