← All GPUs

Apple M3 Max

Apple · Apple Silicon · 2023

400 GB/s

Performance Benchmarks

LLM Inference all sourced rows →

ModelMetricValueSource
Llama-3.1-8B Q4_K_M
Token generation: single user, batch 1, 4k context, mean of 3 runs
tokens per second 55 tok/s
llama.cpp b3500 (Metal) · 2026-05-22
Estimate MyAIHardware llama.cpp Benchmarks ↗
Llama-3.1-8B Q4_K_M
Prompt processing: 512-token prompt, batch 1, mean of 3 runs
prompt tokens per second 1,200 tok/s
llama.cpp b3500 (Metal) · 2026-05-22
Estimate MyAIHardware llama.cpp Benchmarks ↗

Full Specifications

Npu tops
18 TOPS
Gpu cores max
40
Neural engine cores
16
Memory bandwidth gbps
400 GB/s
Unified memory max
128 GB
Architecture
Apple Silicon (M3)

Related GPUs

Guides & Tools

CompareAIHardware.com participates in the Amazon Associates program. As an Amazon Associate, we earn from qualifying purchases.