⌘K
← All AI Models

Can you run IQuest Q1 locally?

IQuest family · 2026 · 320.3B parameters (15B activated per token) · Hugging Face model card

IQuest Q1 is a mixture-of-experts model from IQuest built for agentic coding, reasoning and multi-step tool use, with approximately 320B total parameters and an estimated 15B activated per token (vendor card: Total Parameters 320B; Activated Parameters 15B). The HF safetensors total is 320,318,615,552, consistent with the announced 320B. config.json declares 256 routed experts with 8 selected per token, 88 layers, hidden size 3072, 48 attention heads, 8 KV heads, head dim 128, vocab 160000, and a sliding window of 4096. Context is 524,288 tokens (config.json max_position_embeddings; card lists 524,288). Licensed under the custom IQuest-Q1 license (other). Local VRAM fit uses the 15B active count: roughly 9 GB at Q4-class quantization (calculated). Measured speeds: not yet published here.

Minimum: 10 GB+ memory (Q4-class on 15B active, calculated)  ·  Recommended: 16 GB+ for comfortable context headroom (Q4-class, calculated)

Parameters (total)
320.3B
Activated per token
15B (MoE)
Context window
524,288 tokens
Architecture source
HF config.json (verified)

How much VRAM does IQuest Q1 need at each quantization?

IQuest Q1 needs 212.9 GB of VRAM at Q4_K_M. The table below lists weights-only size and total VRAM including overhead for each common quantization level with 4k and 32k token contexts.

QuantizationBits / weightWeights onlyTotal + KV @4k ctxTotal + KV @32k ctx
Q4_K_M 4.8 192.2 GB 212.9 GB 223.2 GB
Q5_K_M 5.7 228.2 GB 252.5 GB 262.9 GB
Q6_K 6.6 264.3 GB 292.2 GB 302.5 GB
Q8_0 8.5 340.3 GB 375.9 GB 386.2 GB
FP16 16 640.6 GB 706.2 GB 716.5 GB

config.json declares a sliding window of 4096 on attention layers. The card targets agentic coding/tool-use workloads; heavy long-context serving is secondary.

Method: weights = parameters × bits-per-weight (Q4_K_M ≈ 4.8, Q5_K_M ≈ 5.7, Q6_K ≈ 6.6, Q8_0 ≈ 8.5, FP16 = 16), plus 10% loading overhead, plus KV cache from the verified architecture config (88 layers, 8 KV heads, 128 head dim). Source: HF-verified 2026-10-05 via plain api (gated=false safetensors.total=320318615552 params=320.32B ctx=524288 license=other repo=IQuestLab/IQuest-Q1@5c21b0630586ef77d38cff1094b8cf37a417fcd8 layers=88 hidden=3072 kv_heads=8 heads=48 vocab=160000). MoE: num_experts=256 num_experts_per_tok=8 active_params_b=15.0 (vendor model card (Total 320B, Activated 15B per token, stated as approximate/estimated); config num_experts=256 routed 8 per token). VRAM figures are documented calculations (Q4-class ~0.6 GB per B params, on active params for MoE rows), not measurements. release_date is the HF repo lastModified date observed on the verification pass. No tok/s or benchmark scores invented.

Which GPUs can run IQuest Q1 locally?

At Q4_K_M with a 4k context, IQuest Q1 needs 212.9 GB of VRAM. The lists below are computed live from our GPU database and grouped by how much headroom the card has. Disclosure: CompareAIHardware.com participates in the Amazon Associates program and earns from qualifying purchases through links on this page. Affiliate relationships do not influence our recommendations.

Runs IQuest Q1 comfortably (25%+ VRAM headroom)

GPUVRAMBandwidthType
AMD Instinct MI350X 288 GB 8000 GB/s Data Center GPU
AMD Instinct MI355X 288 GB 8000 GB/s Data Center GPU
NVIDIA B300 (Blackwell Ultra) 288 GB 8000 GB/s Data Center GPU
AMD Instinct MI430X 432 GB 23300 GB/s Data Center GPU
AMD Instinct MI455X 432 GB 23300 GB/s Data Center GPU

Minimum GPUs that fit IQuest Q1 at Q4_K_M

These GPUs hold the model but leave little headroom — keep contexts short.

GPUVRAMBandwidthType
AMD Instinct MI325X 256 GB 6000 GB/s Data Center GPU

Needs 2+ GPUs to run IQuest Q1

One of these cards is too small on its own, but a pair (tensor or pipeline parallel, ~90% efficiency) covers the 212.9 GB requirement. See our multi-GPU guide for setup.

GPUVRAMBandwidthType
H200 SXM 141 GB ×2 4800 GB/s Data Center GPU
AMD Instinct MI300X 192 GB ×2 5300 GB/s Data Center GPU
NVIDIA B200 192 GB ×2 8000 GB/s Data Center GPU

Can a Mac run IQuest Q1?

Yes — these Apple Silicon machines fit IQuest Q1 at Q4_K_M, since macOS lets the GPU use about 75% of unified memory. Generation speed is bound by memory bandwidth, so the GB/s column matters as much as capacity.

Apple SiliconUnified memoryUsable by GPU (~75%)Bandwidth
M5 Ultra (Mac Studio) 512 GB 384 GB 1200 GB/s

Frequently asked questions