Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.
Every figure here is computed, not measured. Fit is weights at each quantisation against device memory, with a 25% allowance for the KV cache, activations and the OS. Decode rate is the memory-bandwidth roofline at 70% efficiency. Nobody has run this model on these devices. This is a mixture-of-experts model and no card carries active-parameter counts, so these predictions use total parameters and understate the real speed.
Device
Memory
Bandwidth
Best quality
Weights
tok/s
Fastest
NVIDIA GB200 Grace Blackwell Superchip
372.0 GB
16000.0 GB/s
fp8
237.1 GB
~47.2
~94.5 q4
AMD Instinct MI355X
288.0 GB
8000.0 GB/s
q6
177.82 GB
~31.5
~47.2 q4
NVIDIA B200
180.0 GB
7700.0 GB/s
q4
118.55 GB
~45.5
~45.5 q4
Google TPU7x (Ironwood)
192.0 GB
7380.0 GB/s
q4
118.55 GB
~43.6
~43.6 q4
AMD Instinct MI325X
256.0 GB
6000.0 GB/s
q6
177.82 GB
~23.6
~35.4 q4
AMD Instinct MI300X
192.0 GB
5300.0 GB/s
q4
118.55 GB
~31.3
~31.3 q4
Apple M3 Ultra
512.0 GB
819.0 GB/s
fp8
237.1 GB
~2.4
~4.8 q4
Apple M2 Ultra
192.0 GB
800.0 GB/s
q4
118.55 GB
~4.7
~4.7 q4
Where it runs
Taken from the card's availability section.
Platform
Id on that platform
Hugging Face
LGAI-EXAONE/K-EXAONE-236B-A23B
Verified benchmark evidence
Each row was checked against its source by a reviewer, and carries the model identifier as actually evaluated — which is not always the same as this card's.