NVIDIA DGX Spark vs AMD Ryzen AI Halo (Strix Halo) for Local AI
Updated September 21, 2026. Spec values below come from our product database, pinned to vendor pages. We list measured local speeds only where sourced benchmark rows exist — see the coverage note at the end.
The short answer
Both boxes ship with 128 GB of unified memory at the same $3,999 launch MSRP in our records, so they sit in the same model-size class: large quantized models that do not fit a 24–32 GB discrete card. The DGX Spark is the NVIDIA Grace Blackwell (GB10) personal supercomputer: 273 GB/s LPDDR5x, a 20-core Arm CPU (10 Cortex-X925 + 10 Cortex-A725), CUDA, and NVIDIA’s claim of up to 1 PFLOP at FP4 with sparsity (vendor peak, not a measured tokens/s figure). The Ryzen AI Halo Developer Platform (Max+ 395) is the Strix Halo desktop: 16 Zen 5 cores / 32 threads, Radeon 8060S (40 CU, RDNA 3.5), 256 GB/s LPDDR5X, a 50 TOPS NPU, Windows 11 plus Linux, and a 120 W TDP. Pick Spark when you want the CUDA / DGX OS stack and NVIDIA’s FP4 software path. Pick Halo when you need x86, Windows, ROCm/DirectML, and a slightly lower SoC TDP.
Spec comparison
| Specification | NVIDIA DGX Spark | AMD Ryzen AI Halo (Max+ 395) |
|---|---|---|
| Unified / system memory | 128 GB LPDDR5x unified | 128 GB LPDDR5X 8000 |
| Memory bandwidth | 273 GB/s | 256 GB/s |
| Memory bus | 256-bit | 256-bit |
| CPU | 20-core Arm (10× Cortex-X925 + 10× Cortex-A725) | 16C / 32T Zen 5 |
| GPU | Blackwell integrated in GB10 | Radeon 8060S, 40 CU (RDNA 3.5) |
| NPU | not stated on NVIDIA hardware overview | 50 TOPS |
| Vendor AI peak (estimate / peak, not measured) | up to 1 PFLOP FP4 with sparsity | not listed as PFLOP on the AMD product page we store |
| TDP | 140 W (GB10 SoC); 240 W external PSU | 120 W |
| OS | NVIDIA DGX OS (Ubuntu-based); no Windows | Windows 11, Linux |
| Frameworks (vendor / DB) | CUDA, PyTorch, TRT-LLM | ROCm, Vulkan, DirectML |
| Storage (as listed) | 1 TB or 4 TB NVMe M.2 | 2× NVMe Gen4 SSD |
| Launch MSRP (DB) | $3,999 | $3,999 |
| Release year (DB) | 2025 | 2025 |
Local LLM inference
Token generation on a memory-bound LLM is limited by how fast weights stream from unified memory. Spark’s 273 GB/s versus Halo’s 256 GB/s is a small bandwidth edge on paper (about 7%). Capacity is a tie at 128 GB, so both can hold far larger quantized models than a 16–24 GB card. Software is the larger split: Spark is CUDA-first on Arm + DGX OS; Halo is x86 with ROCm / Vulkan / DirectML and Windows. We do not have sourced head-to-head tokens/s for this pair — treat any speed ranking as unknown until a benchmark row exists.
Image generation and fine-tuning
NVIDIA documents Spark for inference, deployment, and fine-tuning of models up to 200 billion parameters on one unit (405B is listed only for a dual-Spark configuration; that is a vendor topology note, not a measured result). Halo’s 128 GB unified pool and 40 CU Radeon 8060S cover diffusion and smaller fine-tunes on paper; we have no sourced image-gen timings for either box in this comparison.
Which should you buy?
- Buy DGX Spark if you want NVIDIA’s CUDA / TRT-LLM / DGX OS path, FP4 Tensor Core software, and a compact Grace Blackwell desktop with 128 GB unified memory.
- Buy the Ryzen AI Halo Developer Platform if you need Windows or x86 Linux, ROCm/DirectML, a 50 TOPS NPU, and Strix Halo’s 16-core Zen 5 + Radeon 8060S combination at the same listed MSRP.
- Neither number above is a measured tokens/s winner. Bandwidth is close; pick the stack you already run.
How we know (and what we don't)
Every specification above is drawn from our NVIDIA DGX Spark and Ryzen AI Halo Developer Platform product records. Spark numbers are pinned to NVIDIA’s DGX Spark Hardware Overview and the DGX Spark product page. Halo numbers are pinned to the AMD Ryzen AI Max+ 395 / Halo product page stored on that row. Our benchmark database does not yet hold verified head-to-head local-LLM measurements for this pair. Peak PFLOP / TOPS figures are vendor peaks (estimates), not measured throughput. We publish unknowns as unknowns.