⌘K

AI GPU Benchmark Results

Every benchmark number in our database, with its source and evidence type. Filter by workload or evidence type; sort by any column. Independent benchmarks come from third-party reviews and standardized databases (MLPerf); community results are labeled reproducible when the run config is published, estimated when approximated.

100sourced results
27products covered
66GPUs in database
Honest coverage: our current benchmark set covers LLM inference (tokens/s) and Stable Diffusion image generation on flagship and mid-range cards. Gaming FPS and video-generation numbers are not yet sourced — they will appear here when we have verifiable sources. See the methodology and evidence-type definitions.

Filter results

RTX 6000 Ada view all GPUs

100 benchmark results

GPU Workload Model / config Result Evidence Source
RTX 6000 Ada LLM inference (tokens/s) Llama-2-13B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 78.532 tok/s) 3,957 score Independent benchmark0.85 StorageReview GPU Reviews
RTX 6000 Ada LLM inference (tokens/s) Llama-3-70B · Q4 · llama.cpp (48GB VRAM — editorial estimate pending article-level source verification) 35 tok/s EST Estimated0.40 Puget Systems Hardware Testing
RTX 6000 Ada LLM inference (tokens/s) Llama-3-8B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 138.62 tok/s) 4,026 score Independent benchmark0.85 StorageReview GPU Reviews
RTX 6000 Ada LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Token generation: single user, batch 1, 4k context, mean of 3 runs) 110 tok/s Unverified community result0.55 MyAIHardware llama.cpp Benchmarks
RTX 6000 Ada LLM inference (tokens/s) Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (CUDA) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) 3,200 tok/s Unverified community result0.55 MyAIHardware llama.cpp Benchmarks
RTX 6000 Ada LLM inference (tokens/s) Mistral-7B · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 166.633 tok/s) 4,255 score Independent benchmark0.85 StorageReview GPU Reviews
RTX 6000 Ada LLM inference (tokens/s) Phi-3.5-mini · UL Procyon AI Text Generation (TensorRT) (Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 228.359 tok/s) 4,508 score Independent benchmark0.85 StorageReview GPU Reviews
RTX 6000 Ada LLM inference (tokens/s) Qwen3-8B · Q4_K_XL · llama.cpp llama-bench, CUDA 12.8 (Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement) 98.7 tok/s Unverified community result0.55 Hardware Corner LLM GPU Rankings

How these numbers are sourced

Every row links to its original source and carries an evidence type plus a confidence score. Independent benchmarks come from third-party reviews (Tom's Hardware, TechPowerUp, Puget Systems) and standardized databases (MLPerf); manufacturer benchmarks are vendor-published figures; community results are reproducible when the run config is published, unverified otherwise — always labeled EST when estimated. Read the full benchmark methodology.

Compare any two GPUs head-to-head in the comparison tool, or browse value rankings.