⌘K
← All GPUs

GeForce RTX 4090

NVIDIA · Consumer GPU · 2022

24 GB VRAM
1008 GB/s
450W TDP
$1,599 MSRP
Check Latest Price
$1,599 MSRP
Check Price on Amazon →

We earn a commission from qualifying purchases · Prices may vary

Performance Benchmarks

Image Generation all sourced rows →

ModelMetricValueSource
FLUX.1-dev BF16
Requires memory-efficient attention (xFormers/SDPA) within 24GB; approximate
images per minute 4 img/min est
Diffusers · 2026-05-03
Estimated Spheron GPU Benchmark Blog ↗
SDXL Turbo FP16 images per minute 80 img/min
ComfyUI · 2025-06-01
Independent benchmark Tom's Hardware GPU Benchmarks 2025 ↗
Stable Diffusion 1.5 FP16
Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 1.188 s/image
procyon image overall score 5,260 score
UL Procyon AI Image Generation · 2025-01-29
Independent benchmark StorageReview GPU Reviews ↗
Stable Diffusion 1.5 INT8 INT8
Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 0.503 s/image
procyon image overall score 62,160 score
UL Procyon AI Image Generation · 2025-01-29
Independent benchmark StorageReview GPU Reviews ↗
Stable Diffusion XL FP16
Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); 7.461 s/image
procyon image overall score 5,025 score
UL Procyon AI Image Generation · 2025-01-29
Independent benchmark StorageReview GPU Reviews ↗

LLM Inference all sourced rows →

ModelMetricValueSource
Llama-2-13B
Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 92.853 tok/s
procyon text overall score 5,013 score
UL Procyon AI Text Generation (TensorRT) · 2025-01-29
Independent benchmark StorageReview GPU Reviews ↗
Llama-3-70B Q4
24GB VRAM, very tight, heavy swapping — editorial estimate pending article-level source verification
tokens per second 18 tok/s est
llama.cpp · 2025-06-01
Estimated Tom's Hardware GPU Benchmarks 2025 ↗
Llama-3-8B
Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 150.039 tok/s
procyon text overall score 4,849 score
UL Procyon AI Text Generation (TensorRT) · 2025-01-29
Independent benchmark StorageReview GPU Reviews ↗
Llama-3.1-8B FP16
Batch-serving throughput, ~18GB VRAM used; published llama.cpp/vLLM benchmarks (approximate)
batch tokens per second 2,550 tok/s est
vLLM · 2026-05-03
Estimated Spheron GPU Benchmark Blog ↗
Llama-3.1-8B Q4_K_M
Token generation: single user, batch 1, 4k context, mean of 3 runs
tokens per second 125 tok/s
llama.cpp b3500 (CUDA) · 2026-05-22
Unverified community MyAIHardware llama.cpp Benchmarks ↗
Llama-3.1-8B Q4_K_M
Prompt processing: 512-token prompt, batch 1, mean of 3 runs
prompt tokens per second 4,800 tok/s
llama.cpp b3500 (CUDA) · 2026-05-22
Unverified community MyAIHardware llama.cpp Benchmarks ↗
Mistral-7B
Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 183.266 tok/s
procyon text overall score 5,094 score
UL Procyon AI Text Generation (TensorRT) · 2025-01-29
Independent benchmark StorageReview GPU Reviews ↗
Phi-3.5-mini
Standardized UL Procyon suite; same lab platform (ThreadRipper 7980X, driver 571.86); output rate 244.343 tok/s
procyon text overall score 4,958 score
UL Procyon AI Text Generation (TensorRT) · 2025-01-29
Independent benchmark StorageReview GPU Reviews ↗
Qwen3-8B Q4_K_XL
Token generation: 16K context, batch 1, Ubuntu 24.04, on-hardware lab measurement
tokens per second 104.3 tok/s
llama.cpp llama-bench, CUDA 12.8 · 2025-12-09
Unverified community Hardware Corner LLM GPU Rankings ↗

Full Specifications

Vram gb
24 GB
Vram type
GDDR6X
Memory bus
384 bit
Memory bandwidth gbps
1008 GB/s
Tdp watts
450 W
Cuda cores
16384
Pcie interface
PCIe 4.0 x16
Architecture
Ada Lovelace

Related GPUs

Guides & Tools

CompareAIHardware.com participates in the Amazon Associates program. As an Amazon Associate, we earn from qualifying purchases.