← All AI Models

Can you run Stable Diffusion 3.5 Large locally?

Stable Diffusion 3.5 family · October 2024 · 8.00B parameters · Hugging Face model card

Stable Diffusion 3.5 Large runs locally on an 8 GB GPU at Q4, which needs about 7.3 GB of VRAM for the 8-billion-parameter MMDiT weights plus activation overhead. The FP16 model needs about 19.6 GB, so a 24 GB GPU such as the RTX 3090 or RTX 4090 runs SD 3.5 Large at FP16 with the T5 encoder offloaded.

Minimum: 8 GB VRAM GPU at GGUF Q4 with T5 encoder on CPU  ·  Recommended: 24 GB VRAM GPU (RTX 3090/4090) at FP16 or FP8

Parameters (total)
8.00B
VRAM at Q4_K_M (4k ctx)
7.3 GB
Runtime overhead in table
2.0 GB
Architecture source
HF config.json (verified)

How much VRAM does Stable Diffusion 3.5 Large need at each quantization?

Stable Diffusion 3.5 Large needs 7.3 GB of VRAM at Q4_K_M. The table below lists weights-only size and total VRAM including overhead for each common quantization level including runtime overhead.

QuantizationBits / weightWeights onlyTotal + overheadTotal + overhead
Q4_K_M 4.8 4.8 GB 7.3 GB 7.3 GB
Q5_K_M 5.7 5.7 GB 8.3 GB 8.3 GB
Q6_K 6.6 6.6 GB 9.3 GB 9.3 GB
Q8_0 8.5 8.5 GB 11.4 GB 11.4 GB
FP16 16 16.0 GB 19.6 GB 19.6 GB

MMDiT-X diffusion transformer — no KV cache. Table adds a 2 GB activation overhead; the T5-XXL text encoder (4.7B params, ~9.4 GB FP16) can be offloaded to CPU RAM or swapped for smaller encoders.

Method: weights = parameters × bits-per-weight (Q4_K_M ≈ 4.8, Q5_K_M ≈ 5.7, Q6_K ≈ 6.6, Q8_0 ≈ 8.5, FP16 = 16), plus 10% loading overhead, plus KV cache from the verified architecture config (not applicable to this architecture). Source: 8B-parameter MMDiT model; SD 3.5 Medium is 2.5B — Stability AI / HF model card (stabilityai/stable-diffusion-3.5-large).

Which GPUs can run Stable Diffusion 3.5 Large locally?

At Q4_K_M with a 4k context, Stable Diffusion 3.5 Large needs 7.3 GB of VRAM. The lists below are computed live from our GPU database and grouped by how much headroom the card has. Disclosure: CompareAIHardware.com participates in the Amazon Associates program and earns from qualifying purchases through links on this page. Affiliate relationships do not influence our recommendations.

Runs Stable Diffusion 3.5 Large comfortably (25%+ VRAM headroom)

GPUVRAMBandwidthType
Arc B570 10 GB 380 GB/s Consumer GPU Check price
GeForce RTX 3080 10GBEOL 10 GB 760 GB/s Consumer GPU Check price
GeForce RTX 2080 TiEOL 11 GB 616 GB/s Consumer GPU Check price
Arc B580 12 GB 456 GB/s Consumer GPU Check price
GeForce RTX 3060 12GB 12 GB 360 GB/s Consumer GPU Check price
GeForce RTX 3080 TiEOL 12 GB 912 GB/s Consumer GPU Check price
GeForce RTX 4070EOL 12 GB 504 GB/s Consumer GPU Check price
GeForce RTX 4070 SUPEREOL 12 GB 504 GB/s Consumer GPU Check price
GeForce RTX 5070 12 GB 672 GB/s Consumer GPU Check price
Radeon RX 7700 XT 12 GB 432 GB/s Consumer GPU Check price
Arc A770 16GB 16 GB 560 GB/s Consumer GPU Check price
GeForce RTX 4060 Ti 16GB 16 GB 288 GB/s Consumer GPU Check price
GeForce RTX 4070 Ti SUPEREOL 16 GB 672 GB/s Consumer GPU Check price
GeForce RTX 4080EOL 16 GB 716 GB/s Consumer GPU Check price
GeForce RTX 4080 SUPER 16 GB 736 GB/s Consumer GPU Check price
GeForce RTX 5060 Ti 16GB 16 GB 448 GB/s Consumer GPU Check price
GeForce RTX 5070 Ti 16 GB 896 GB/s Consumer GPU Check price
GeForce RTX 5080 16 GB 960 GB/s Consumer GPU Check price
NVIDIA RTX PRO 2000 Blackwell 16 GB 288 GB/s Pro GPU Check price
Radeon RX 6800 XTEOL 16 GB 512 GB/s Consumer GPU Check price
Radeon RX 7600 XT 16 GB 288 GB/s Consumer GPU Check price
Radeon RX 7800 XT 16 GB 624 GB/s Consumer GPU Check price
Radeon RX 9070 16 GB 640 GB/s Consumer GPU Check price
Radeon RX 9070 XT 16 GB 640 GB/s Consumer GPU Check price
Radeon RX 7900 XT 20 GB 800 GB/s Consumer GPU Check price
GeForce RTX 3090EOL 24 GB 936 GB/s Consumer GPU Check price
GeForce RTX 3090 TiEOL 24 GB 1008 GB/s Consumer GPU Check price
GeForce RTX 4090 24 GB 1008 GB/s Consumer GPU Check price
NVIDIA RTX PRO 4000 Blackwell 24 GB 672 GB/s Pro GPU Check price
Radeon RX 7900 XTX 24 GB 960 GB/s Consumer GPU Check price
AMD Radeon Pro W7800 32 GB 576 GB/s Pro GPU Check price
GeForce RTX 5090 32 GB 1792 GB/s Consumer GPU Check price
NVIDIA RTX 5000 Ada 32 GB 576 GB/s Pro GPU Check price
NVIDIA RTX PRO 4500 Blackwell 32 GB 896 GB/s Pro GPU Check price
NVIDIA A100 40GB SXMEOL 40 GB 1555 GB/s Data Center GPU
AMD Radeon Pro W7900 48 GB 864 GB/s Pro GPU Check price
NVIDIA RTX PRO 5000 Blackwell 48 GB 1344 GB/s Pro GPU
RTX 6000 Ada 48 GB 960 GB/s Pro GPU Check price
RTX A6000 48 GB 768 GB/s Pro GPU Check price
H100 SXM 80 GB 3350 GB/s Data Center GPU
NVIDIA A100 80GB SXM 80 GB 2039 GB/s Data Center GPU
NVIDIA H100 PCIe 80GB 80 GB 2039 GB/s Data Center GPU
NVIDIA RTX PRO 6000 Blackwell 96 GB 1792 GB/s Pro GPU Check price
H200 SXM 141 GB 4800 GB/s Data Center GPU
AMD Instinct MI300X 192 GB 5300 GB/s Data Center GPU
NVIDIA B200 192 GB 8000 GB/s Data Center GPU
AMD Instinct MI325X 256 GB 6000 GB/s Data Center GPU
AMD Instinct MI355X 288 GB 8000 GB/s Data Center GPU
NVIDIA B300 (Blackwell Ultra) 288 GB 8000 GB/s Data Center GPU

Minimum GPUs that fit Stable Diffusion 3.5 Large at Q4_K_M

These GPUs hold the model but leave little headroom — keep contexts short.

GPUVRAMBandwidthType
Arc A580 8 GB 512 GB/s Consumer GPU Check price
Arc A750 8 GB 512 GB/s Consumer GPU Check price
GeForce RTX 3070EOL 8 GB 448 GB/s Consumer GPU Check price
GeForce RTX 4060 8 GB 272 GB/s Consumer GPU Check price
GeForce RTX 4060 Ti 8GB 8 GB 288 GB/s Consumer GPU Check price
GeForce RTX 5050 8 GB 320 GB/s Consumer GPU Check price
GeForce RTX 5060 8 GB 448 GB/s Consumer GPU Check price
Radeon RX 7600 8 GB 288 GB/s Consumer GPU Check price

Can a Mac run Stable Diffusion 3.5 Large?

Yes — these Apple Silicon machines fit Stable Diffusion 3.5 Large at Q4_K_M, since macOS lets the GPU use about 75% of unified memory. Generation speed is bound by memory bandwidth, so the GB/s column matters as much as capacity.

Apple SiliconUnified memoryUsable by GPU (~75%)Bandwidth
Apple M3 24 GB 18 GB 150 GB/s
Apple M4 32 GB 24 GB 120 GB/s
Apple M3 Pro 36 GB 27 GB 300 GB/s
Apple M4 Pro 64 GB 48 GB 273 GB/s
Apple M3 Max 128 GB 96 GB 400 GB/s
M4 Max (MacBook Pro) 128 GB 96 GB 546 GB/s
M3 Ultra (Mac Studio) 512 GB 384 GB 819 GB/s

Can Stable Diffusion 3.5 Large run on mini PCs, Jetson, or NPU devices?

These edge and NPU devices from our database have enough memory for Stable Diffusion 3.5 Large at Q4_K_M. Their memory bandwidth is far below discrete GPUs, so expect a fraction of desktop generation speed.

DeviceMemoryBandwidthType
Jetson Orin Nano 8GB Super (Dev Kit) 8 GB 102 GB/s Edge Compute Device
Jetson Orin NX 16GB Super 16 GB 102 GB/s Edge Compute Device
Intel Core Ultra 9 288V (Lunar Lake) 32 GB 136 GB/s NPU Chip
Jetson AGX Orin 32GB 32 GB 204 GB/s Edge Compute Device
Jetson AGX Orin 64GB 64 GB 204 GB/s Edge Compute Device
Snapdragon X Elite (X1E-84-100) 64 GB 135 GB/s NPU Chip
AMD Ryzen AI 9 HX 370 (Strix Point) 96 GB 120 GB/s NPU Chip

Which pre-built systems can run Stable Diffusion 3.5 Large?

These mini PCs and workstations from our database fit Stable Diffusion 3.5 Large at Q4_K_M. Usable-memory figures are conservative: Windows shares about half of system RAM with the GPU by default (Linux can expose more), macOS lets Apple Silicon GPUs use about 75% of unified memory, and Linux unified-memory systems such as GB10 expose roughly 90%. Generation speed is bound by memory bandwidth, so compare the GB/s column before buying.

SystemTypeMemoryUsable for AIBandwidthGPUs
GMKtec NucBox 10 (Ryzen 7 5800U) Mini PC 16 GB 8 GB 51 GB/s
Mac mini M4 (Base) Mini PC 16 GB 8 GB 120 GB/s
ASUS NUC 14 Pro+ (Core Ultra 7 258V) Mini PC 32 GB 16 GB 136 GB/s
Beelink SER6 MAX (Ryzen 9 6900HX) Mini PC 32 GB 16 GB 76 GB/s
Beelink SER7 (Ryzen 7 7840HS) Mini PC 32 GB 16 GB 83 GB/s
Beelink SER8 (Ryzen 7 8845HS) Mini PC 32 GB 16 GB 89 GB/s
Beelink SER9 (Ryzen AI 9 HX 370) Mini PC 32 GB 16 GB 120 GB/s
Lenovo ThinkCentre Neo 50q (Core Ultra 7 258V) Mini PC 32 GB 16 GB 136 GB/s
Lenovo ThinkCentre Tiny (Core Ultra 5 125U) Mini PC 32 GB 16 GB 119 GB/s
MSI Cubi 5 (Core i5-13420H) Mini PC 32 GB 16 GB 83 GB/s
MSI Cubi NUC AI+ (Core Ultra 7 258V) Mini PC 32 GB 16 GB 136 GB/s
MinisForum EM680 (Ryzen 7 6800H) Mini PC 32 GB 16 GB 76 GB/s
MinisForum UM890 Slim (Ryzen 7 8845HS) Mini PC 32 GB 16 GB 89 GB/s
Mac mini M4 Max Mini PC 48 GB 24 GB 546 GB/s
Mac mini M4 Pro Mini PC 48 GB 24 GB 273 GB/s
ASRock DeskMeet X600 (Ryzen 7 7600) Mini PC 64 GB 32 GB 83 GB/s
BIZON G3000 G2 (1x RTX 5090) Workstation 32 GB 32 GB 1792 GB/s 1
Beelink SER9 Pro (Ryzen AI 9 HX 370, 64GB) Mini PC 64 GB 32 GB 120 GB/s
Custom RTX 5090 AI Workstation (Value Build) Workstation 32 GB 32 GB 1792 GB/s 1
HP Elite Mini 800 G9 (Core i7-13700T) Mini PC 64 GB 32 GB 83 GB/s
HP OMEN 45L RTX 5090 Desktop Workstation 32 GB 32 GB 1
Intel NUC 13 Pro (Core i7-1360P) Mini PC 64 GB 32 GB 83 GB/s
MinisForum UM780 XTX (Ryzen 7 7840HS) Mini PC 64 GB 32 GB 89 GB/s
MinisForum UM890 Pro (Ryzen 9 8945HS) Mini PC 64 GB 32 GB 89 GB/s
Puget Systems Datum (1x RTX 5090) Workstation 32 GB 32 GB 1792 GB/s 1
Sentinel RTX 5090 Tower Workstation Workstation 32 GB 32 GB 1
ASRock NUC BOX-255H (Core Ultra 7 255H) Mini PC 96 GB 48 GB 102 GB/s
Apple Mac Studio M2 Ultra (64GB) Workstation 64 GB 48 GB 800 GB/s 1
Custom Dual RTX 4090 Training Workstation Workstation 48 GB 48 GB 1008 GB/s 2
Dell Precision 7960 Tower (1x RTX 6000 Ada) Workstation 48 GB 48 GB 960 GB/s 1
HP Z8 Fury G5 (1x RTX 6000 Ada) Workstation 48 GB 48 GB 960 GB/s 1
AMD Ryzen AI Halo Developer Platform (Max+ 395) Mini PC 128 GB 64 GB 256 GB/s
ArsenalPC MES2X Dual RTX 5090 AI Workstation Workstation 64 GB 64 GB 2
GMKtec EVO-X2 (Ryzen AI Max+ 395) Mini PC 128 GB 64 GB 256 GB/s
Mac mini M4 Max (128GB) Mini PC 128 GB 64 GB 546 GB/s
MinisForum MS-S1 MAX (Ryzen AI Max+ 395) Mini PC 128 GB 64 GB 256 GB/s
BOXX APEXX 8R (1x RTX PRO 6000 Blackwell) Workstation 96 GB 96 GB 1792 GB/s 1
GEEKOM A9 Mega AI Workstation Workstation 128 GB 96 GB 1
NOVATECH RTX PRO 6000 AI Workstation Workstation 96 GB 96 GB 1
System76 Thelio Major (1x RTX PRO 6000 Blackwell) Workstation 96 GB 96 GB 1792 GB/s 1
System76 Thelio Major (2x RTX 6000 Ada) Workstation 96 GB 96 GB 960 GB/s 2
Lenovo ThinkStation PGX Workstation 128 GB 115.2 GB 273 GB/s 1
NVIDIA DGX Spark Workstation 128 GB 115.2 GB 273 GB/s 1
BIZON G3000 G2 (4x RTX 5090) Workstation 128 GB 128 GB 1792 GB/s 4
Apple Mac Studio M2 Ultra (192GB) Workstation 192 GB 144 GB 800 GB/s 1
HP Z8 Fury G5 (4x RTX 6000 Ada) Workstation 192 GB 192 GB 960 GB/s 4

Frequently asked questions

How much VRAM does Stable Diffusion 3.5 Large need?

Stable Diffusion 3.5 Large needs about 7.3 GB of VRAM at Q4 and about 19.6 GB at FP16 for the diffusion model alone, plus up to 9.4 GB for the T5-XXL text encoder at FP16 unless you offload it to system RAM.

Can an RTX 3060 12GB run SD 3.5 Large?

Yes, at Q4 or FP8 quantization with the T5 encoder offloaded to CPU, which needs about 7.3–11 GB of VRAM. FP16 does not fit 12 GB, so expect slower first-token latency from the offloaded encoder.

What is the difference between SD 3.5 Large and Medium?

SD 3.5 Large is the 8-billion-parameter flagship for prompt adherence and text rendering, while SD 3.5 Medium is a 2.5-billion-parameter model tuned for consumer GPUs between 8 GB and 10 GB of VRAM, according to Stability AI’s model cards.

Related guides