← All GPUs
GeForce RTX 5090
NVIDIA · Consumer GPU · 2025
Check Latest Price
$1,999 MSRP
Check Price on Amazon →
We earn a commission from qualifying purchases · Prices may vary
Performance Benchmarks
Image Generation all sourced rows →
| Model | Metric | Value | Source |
|---|---|---|---|
|
FLUX.1-dev
BF16 ~26GB VRAM used; community + Spheron internal testing (approximate) |
images per minute |
5.5
img/min
est Diffusers · 2026-05-03
|
Estimated Spheron GPU Benchmark Blog ↗ |
| SDXL Turbo FP16 | images per minute |
120
img/min
ComfyUI · 2025-06-01
|
Independent benchmark Tom's Hardware GPU Benchmarks 2025 ↗ |
|
Stable Diffusion 1.5
FP16 Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 0.763 s/image |
procyon image overall score |
8,193
score
UL Procyon AI Image Generation · 2025-01-29
|
Independent benchmark StorageReview GPU Reviews ↗ |
|
Stable Diffusion 1.5 INT8
INT8 Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 0.394 s/image |
procyon image overall score |
79,272
score
UL Procyon AI Image Generation · 2025-01-29
|
Independent benchmark StorageReview GPU Reviews ↗ |
|
Stable Diffusion XL
FP16 Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 5.223 s/image |
procyon image overall score |
7,179
score
UL Procyon AI Image Generation · 2025-01-29
|
Independent benchmark StorageReview GPU Reviews ↗ |
LLM Inference all sourced rows →
| Model | Metric | Value | Source |
|---|---|---|---|
|
Llama-2-13B
Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 134.502 tok/s |
procyon text overall score |
6,591
score
UL Procyon AI Text Generation (TensorRT) · 2025-01-29
|
Independent benchmark StorageReview GPU Reviews ↗ |
|
Llama-3-70B
Q4 32GB VRAM, tight fit — editorial estimate pending article-level source verification |
tokens per second |
42
tok/s
est llama.cpp · 2025-06-01
|
Estimated Tom's Hardware GPU Benchmarks 2025 ↗ |
|
Llama-3-8B
Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 214.285 tok/s |
procyon text overall score |
6,104
score
UL Procyon AI Text Generation (TensorRT) · 2025-01-29
|
Independent benchmark StorageReview GPU Reviews ↗ |
|
Llama-3.1-8B
FP16 Batch-serving throughput, ~18GB VRAM used; community + Spheron internal testing (approximate) |
batch tokens per second |
3,500
tok/s
est vLLM · 2026-05-03
|
Estimated Spheron GPU Benchmark Blog ↗ |
|
Llama-3.1-8B
Q4_K_M Token generation: single user, batch 1, 4k context, mean of 3 runs |
tokens per second |
215
tok/s
llama.cpp b3500 (CUDA) · 2026-05-22
|
Unverified community MyAIHardware llama.cpp Benchmarks ↗ |
|
Llama-3.1-8B
Q4_K_M Prompt processing: 512-token prompt, batch 1, mean of 3 runs |
prompt tokens per second |
9,800
tok/s
llama.cpp b3500 (CUDA) · 2026-05-22
|
Unverified community MyAIHardware llama.cpp Benchmarks ↗ |
|
Mistral-7B
Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 255.945 tok/s |
procyon text overall score |
6,267
score
UL Procyon AI Text Generation (TensorRT) · 2025-01-29
|
Independent benchmark StorageReview GPU Reviews ↗ |
|
Phi-3.5-mini
Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 314.435 tok/s |
procyon text overall score |
5,749
score
UL Procyon AI Text Generation (TensorRT) · 2025-01-29
|
Independent benchmark StorageReview GPU Reviews ↗ |
|
Qwen3-8B
Q4_K_XL Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement |
tokens per second |
145.3
tok/s
llama.cpp llama-bench, CUDA 12.8 · 2025-12-09
|
Unverified community Hardware Corner LLM GPU Rankings ↗ |
Full Specifications
Vram gb
32 GB
Vram type
GDDR7
Memory bus
512 bit
Memory bandwidth gbps
1792 GB/s
Tdp watts
575 W
Cuda cores
21760
Pcie interface
PCIe 5.0 x16
Architecture
Blackwell
Related GPUs
Guides & Tools
CompareAIHardware.com participates in the Amazon Associates program. As an Amazon Associate, we earn from qualifying purchases.