AI GPU Benchmark Results
Every benchmark number in our database, with its source and tier. Filter by workload or source tier; sort by any column. Tier 1 = manufacturer/benchmark-DB, Tier 2 = independent reviews, Tier 3 = community or estimated.
97sourced results
26products covered
64GPUs in database
Honest coverage: our current benchmark set covers LLM inference (tokens/s) and Stable Diffusion image generation on flagship and mid-range cards. Gaming FPS and video-generation numbers are not yet sourced — they will appear here when we have verifiable sources. See the methodology and tier definitions.
Filter results
97 benchmark results
| GPU | Workload | Model / config | Result | Source tier | Source |
|---|---|---|---|---|---|
| Apple M3 Max | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (Metal) (Token generation: single user, batch 1, 4k context, mean of 3 runs) | 55 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
| Apple M3 Max | LLM inference (tokens/s) | Llama-3.1-8B · Q4_K_M · llama.cpp b3500 (Metal) (Prompt processing: 512-token prompt, batch 1, mean of 3 runs) | 1,200 tok/s | T3 | MyAIHardware llama.cpp Benchmarks |
How these numbers are sourced
Every row links to its original source. Tier 1 comes from manufacturers or standardized benchmark databases (MLPerf), Tier 2 from independent reviews (Tom's Hardware, TechPowerUp, Puget Systems), Tier 3 from community measurements or our own scaling estimates — always labeled EST when estimated. Read the full benchmark methodology.
Compare any two GPUs head-to-head in the comparison tool, or browse value rankings.