⌘K
← All AI Models

Can you run Nex N2.5 Max locally?

Nex family · Sep 2026 · 1.6T parameters · Hugging Face model card

Nex N2.5 Max is a mixture-of-experts model activating 6 of 384 routed experts per token (plus 1 shared) released with open weights on Hugging Face: 1600.79B total parameters, a 1,048,576-token context window, apache-2.0 license. Running it locally needs roughly 1024 GB of memory at Q4-class quantization (calculated) — multi-GPU server or large-unified-memory territory. Measured local speeds: not yet published here; we list unknowns as unknowns.

Minimum: 1024 GB+ memory (Q4-class, calculated)  ·  Recommended: 1536 GB+ for comfortable context headroom (Q4-class, calculated)

Parameters (total)
1.6T
VRAM at Q4_K_M (4k ctx)
1,056.5 GB
Context window
1,048,576 tokens
Architecture source
HF config.json (verified)

How much VRAM does Nex N2.5 Max need at each quantization?

Nex N2.5 Max needs 1,056.5 GB of VRAM at Q4_K_M. The table below lists weights-only size and total VRAM including overhead for each common quantization level with 4k and 32k token contexts.

QuantizationBits / weightWeights onlyTotal + KV @4k ctxTotal + KV @32k ctx
Q4_K_M 4.8 960.5 GB 1,056.5 GB 1,056.5 GB
Q5_K_M 5.7 1,140.6 GB 1,254.6 GB 1,254.6 GB
Q6_K 6.6 1,320.7 GB 1,452.7 GB 1,452.7 GB
Q8_0 8.5 1,700.8 GB 1,870.9 GB 1,870.9 GB
FP16 16 3,201.6 GB 3,521.7 GB 3,521.7 GB

Method: weights = parameters × bits-per-weight (Q4_K_M ≈ 4.8, Q5_K_M ≈ 5.7, Q6_K ≈ 6.6, Q8_0 ≈ 8.5, FP16 = 16), plus 10% loading overhead, plus KV cache from the verified architecture config (not applicable to this architecture). Source: HF-verified 2026-09-16 run2 via api (safetensors + config.json: params=1600.79B ctx=1,048,576 license=apache-2.0); VRAM figures are documented calculations, not measurements; active params: config.json n_routed_experts=384 num_experts_per_tok=6 n_shared_experts=1; per-token active param count not published in config — left null

Which GPUs can run Nex N2.5 Max locally?

At Q4_K_M with a 4k context, Nex N2.5 Max needs 1,056.5 GB of VRAM. The lists below are computed live from our GPU database and grouped by how much headroom the card has. Disclosure: CompareAIHardware.com participates in the Amazon Associates program and earns from qualifying purchases through links on this page. Affiliate relationships do not influence our recommendations.

No single consumer GPU in our database runs Nex N2.5 Max at Q4_K_M. The 1,056.5 GB requirement exceeds every listed consumer card; see the Mac and data center options above and below.

Frequently asked questions