Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.
Every figure here is computed, not measured. Fit is weights at each quantisation against device memory, with a 25% allowance for the KV cache, activations and the OS. Decode rate is the memory-bandwidth roofline at 70% efficiency. Nobody has run this model on these devices.
Device
Memory
Bandwidth
Best quality
Weights
tok/s
Fastest
NVIDIA GB200 Grace Blackwell Superchip
372.0 GB
16000.0 GB/s
bf16
205.78 GB
~910.3
~3641.1 q4
AMD Instinct MI355X
288.0 GB
8000.0 GB/s
bf16
205.78 GB
~455.1
~1820.5 q4
NVIDIA B200
180.0 GB
7700.0 GB/s
fp8
102.89 GB
~876.1
~1752.3 q4
Google TPU7x (Ironwood)
192.0 GB
7380.0 GB/s
fp8
102.89 GB
~839.7
~1679.5 q4
AMD Instinct MI325X
256.0 GB
6000.0 GB/s
fp8
102.89 GB
~682.7
~1365.4 q4
AMD Instinct MI300X
192.0 GB
5300.0 GB/s
fp8
102.89 GB
~603.1
~1206.1 q4
NVIDIA H200 SXM
141.0 GB
4800.0 GB/s
fp8
102.89 GB
~546.2
~1092.3 q4
NVIDIA H100 NVL
94.0 GB
3900.0 GB/s
q5
64.31 GB
~710.0
~887.5 q4
NVIDIA H100 SXM
80.0 GB
3350.0 GB/s
q4
51.44 GB
~762.4
~762.4 q4
AMD Instinct MI250X
128.0 GB
3200.0 GB/s
q6
77.17 GB
~485.5
~728.2 q4
Google TPU v5p
95.0 GB
2765.0 GB/s
q5
64.31 GB
~503.4
~629.2 q4
NVIDIA A100 80GB SXM
80.0 GB
2039.0 GB/s
q4
51.44 GB
~464.0
~464.0 q4
NVIDIA H100 PCIe
80.0 GB
2000.0 GB/s
q4
51.44 GB
~455.1
~455.1 q4
Apple M3 Ultra
512.0 GB
819.0 GB/s
bf16
205.78 GB
~46.6
~186.4 q4
Apple M2 Ultra
192.0 GB
800.0 GB/s
fp8
102.89 GB
~91.0
~182.1 q4
Apple M4 Max
128.0 GB
546.0 GB/s
q6
77.17 GB
~82.8
~124.3 q4
Apple M2 Max
96.0 GB
400.0 GB/s
q5
64.31 GB
~72.8
~91.0 q4
Apple M3 Max
128.0 GB
400.0 GB/s
q6
77.17 GB
~60.7
~91.0 q4
Where it runs
Taken from the card's availability section.
Platform
Id on that platform
Hugging Face
inclusionAI/Ling-flash-2.0
Verified benchmark evidence
Each row was checked against its source by a reviewer, and carries the model identifier as actually evaluated — which is not always the same as this card's.