← All GPUs
Apple M3 Max
Apple · Apple Silicon · 2023
Performance Benchmarks
LLM Inference all sourced rows →
| Model | Metric | Value | Source |
|---|---|---|---|
|
Llama-3.1-8B
Q4_K_M Token generation: single user, batch 1, 4k context, mean of 3 runs |
tokens per second |
55
tok/s
llama.cpp b3500 (Metal) · 2026-05-22
|
Estimate MyAIHardware llama.cpp Benchmarks ↗ |
|
Llama-3.1-8B
Q4_K_M Prompt processing: 512-token prompt, batch 1, mean of 3 runs |
prompt tokens per second |
1,200
tok/s
llama.cpp b3500 (Metal) · 2026-05-22
|
Estimate MyAIHardware llama.cpp Benchmarks ↗ |
Full Specifications
Npu tops
18 TOPS
Gpu cores max
40
Neural engine cores
16
Memory bandwidth gbps
400 GB/s
Unified memory max
128 GB
Architecture
Apple Silicon (M3)
Related GPUs
Guides & Tools
CompareAIHardware.com participates in the Amazon Associates program. As an Amazon Associate, we earn from qualifying purchases.