⌘K

NVIDIA DGX Spark vs Apple M3 Ultra 96GB for Local AI

Updated September 21, 2026. Spec values below come from our product database, pinned to vendor pages. We list measured local speeds only where sourced benchmark rows exist — see the coverage note at the end.

The short answer

These two desktops trade capacity versus bandwidth. DGX Spark has 128 GB of LPDDR5x unified memory at 273 GB/s on a Grace Blackwell GB10 Superchip (20-core Arm, integrated Blackwell GPU, CUDA, DGX OS). The M3 Ultra Mac Studio 96 GB has 96 GB unified memory but 819 GB/s — about 3× Spark’s bandwidth — plus up to 80 GPU cores and a 32-core Neural Engine, on macOS / Metal. Both rows list a $3,999 launch MSRP. Spark wins raw resident model size. M3 Ultra wins memory bandwidth on the stored specs. Software is CUDA versus Metal; neither side has a sourced tokens/s row in our benchmark table for this pair.

Spec comparison

SpecificationNVIDIA DGX SparkM3 Ultra (Mac Studio, 96 GB)
Unified memory128 GB LPDDR5x96 GB (unified_memory_max)
Memory bandwidth273 GB/s819 GB/s
CPU / SoCGB10: 20-core Arm (10× Cortex-X925 + 10× Cortex-A725)Apple Silicon (M3 Ultra)
GPUBlackwell integrated in GB10up to 80 GPU cores
NPU / Neural Enginenot stated on NVIDIA hardware overview32 cores
Vendor AI peak (estimate / peak, not measured)up to 1 PFLOP FP4 with sparsitynot stored as PFLOP on this product row
TDP140 W GB10 SoC (240 W external PSU)270 W
OS / APIDGX OS (Ubuntu-based), CUDA; no Windows, no MetalmacOS, Metal
Launch MSRP (DB)$3,999$3,999
Release year (DB)20252025

Local LLM inference

Spark can keep a larger quantized model resident (128 GB vs 96 GB). M3 Ultra should generate tokens faster on bandwidth-bound decode if software can feed the GPU, because 819 GB/s is about 3× 273 GB/s. That is a spec-ratio statement, not a measured speedup. Tooling splits the market: llama.cpp / MLX / Metal on Apple versus CUDA / TRT-LLM / PyTorch on Spark. We have no sourced head-to-head tokens/s for this pair.

Image generation and fine-tuning

NVIDIA lists Spark for inference, deployment, and fine-tuning up to 200B parameters on one unit. Apple’s stored row does not publish an equivalent parameter-count claim; 96 GB still covers large quantized diffusion and mid-size fine-tunes on paper. No sourced image-gen timings for either product appear in this article.

Which should you buy?

  • Buy DGX Spark if you need 128 GB unified memory, CUDA, and NVIDIA’s DGX software stack on a compact GB10 desktop.
  • Buy the M3 Ultra 96 GB Mac Studio if you already live in macOS / Metal / MLX and want the much higher listed memory bandwidth, accepting 32 GB less unified capacity.
  • Need more than 96 GB on Apple? Our M3 Ultra 256 GB row exists as a higher-capacity SKU; this page compares the 96 GB SKU only.

How we know (and what we don't)

Every specification above is drawn from our NVIDIA DGX Spark and M3 Ultra 96 GB product records. Spark numbers are pinned to NVIDIA’s DGX Spark Hardware Overview and the DGX Spark product page. M3 Ultra numbers are pinned to Apple support document 111901 via the product source_url. Peak PFLOP is a vendor peak (estimate), not measured throughput. Our benchmark database does not yet hold verified head-to-head local-LLM measurements for this pair. We publish unknowns as unknowns.