Benchmark Methodology
Last updated: July 20, 2026
Core Principle
Every benchmark has provenance. Every recommendation has reasoning. We distinguish tested results from estimates — always.
Source Classification
Tier 1: Verified Primary Sources
- Manufacturer official specifications (NVIDIA, AMD, Intel)
- Our own controlled testing
- Published academic papers with reproducible methodology
Tier 2: Trusted Secondary Sources
- Established tech publications (Tom's Hardware, TechPowerUp, etc.)
- Community-maintained databases with editorial oversight (MLPerf)
- Developer-reported results from model creators (Meta, Mistral, DeepSeek)
Tier 3: Community Data
- Reddit/forum benchmarks with sufficient detail
- User-submitted results (labeled as unverified)
What We Measure
LLM Inference
- Tokens per second (t/s): Generation throughput at specified quantization
- Time to first token (TTFT): Prompt processing latency
- Context length tested: Always reported alongside results
- Model and quantization: e.g., Llama-3-70B Q4_K_M
- Software stack: llama.cpp, vLLM, Ollama version noted
Image Generation
- Time to generate: For specific resolutions and models (SDXL, SD3, Flux)
- Batch size and steps: Always noted
Training
- Training throughput: Samples/sec or iterations/sec
- Model size and batch size: Always noted
- Mixed precision: FP16, BF16, FP8 noted
Value Metrics
- Performance per dollar: Benchmark result divided by current street price
- Performance per watt: Benchmark result divided by TDP or measured power
- Performance per VRAM-GB: Comparing efficiency across memory tiers
All prices are USD with date captured. Used-market prices noted separately from MSRP.
Estimates vs. Tested
When tested data is unavailable, we may provide estimates based on architectural scaling, similar GPU results, or model requirements. Estimates are always clearly labeled with reasoning shown.
Data Currency
Each benchmark data point includes date tested, software version, and source link. We re-test when major updates occur.