Moonshot AI · moonshot/kimi-linear-48b-a3b-instruct
Provider
Moonshot AI
Family
kimi
Type
llm-chat
Status
active
Released
2025-10-30
Updated
2025-12-16
Parameters
49.1B
Open weights
yes
What it fits on
Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.
Every figure here is computed, not measured. Fit is weights at each quantisation against device memory, with a 25% allowance for the KV cache, activations and the OS. Decode rate is the memory-bandwidth roofline at 70% efficiency. Nobody has run this model on these devices. This is a mixture-of-experts model and no card carries active-parameter counts, so these predictions use total parameters and understate the real speed.
Device
Memory
Bandwidth
Best quality
Weights
tok/s
Fastest
Cerebras WSE-3
44.0 GB
21000000.0 GB/s
q5
30.7 GB
~478801.2
~598501.5 q4
NVIDIA GB200 Grace Blackwell Superchip
372.0 GB
16000.0 GB/s
bf16
98.25 GB
~114.0
~456.0 q4
AMD Instinct MI355X
288.0 GB
8000.0 GB/s
bf16
98.25 GB
~57.0
~228.0 q4
NVIDIA B200
180.0 GB
7700.0 GB/s
bf16
98.25 GB
~54.9
~219.5 q4
Google TPU7x (Ironwood)
192.0 GB
7380.0 GB/s
bf16
98.25 GB
~52.6
~210.3 q4
AMD Instinct MI325X
256.0 GB
6000.0 GB/s
bf16
98.25 GB
~42.8
~171.0 q4
AMD Instinct MI300X
192.0 GB
5300.0 GB/s
bf16
98.25 GB
~37.8
~151.1 q4
NVIDIA H200 SXM
141.0 GB
4800.0 GB/s
bf16
98.25 GB
~34.2
~136.8 q4
NVIDIA H100 NVL
94.0 GB
3900.0 GB/s
fp8
49.12 GB
~55.6
~111.2 q4
NVIDIA H100 SXM
80.0 GB
3350.0 GB/s
fp8
49.12 GB
~47.7
~95.5 q4
AMD Instinct MI250X
128.0 GB
3200.0 GB/s
fp8
49.12 GB
~45.6
~91.2 q4
Google TPU v5p
95.0 GB
2765.0 GB/s
fp8
49.12 GB
~39.4
~78.8 q4
NVIDIA A100 80GB SXM
80.0 GB
2039.0 GB/s
fp8
49.12 GB
~29.1
~58.1 q4
NVIDIA H100 PCIe
80.0 GB
2000.0 GB/s
fp8
49.12 GB
~28.5
~57.0 q4
AMD Instinct MI210
64.0 GB
1600.0 GB/s
q6
36.84 GB
~30.4
~45.6 q4
NVIDIA A100 40GB SXM
40.0 GB
1555.0 GB/s
q4
24.56 GB
~44.3
~44.3 q4
NVIDIA L40S
48.0 GB
864.0 GB/s
q5
30.7 GB
~19.7
~24.6 q4
Apple M3 Ultra
512.0 GB
819.0 GB/s
bf16
98.25 GB
~5.8
~23.3 q4
Apple M2 Ultra
192.0 GB
800.0 GB/s
bf16
98.25 GB
~5.7
~22.8 q4
Apple M4 Max
128.0 GB
546.0 GB/s
fp8
49.12 GB
~7.8
~15.6 q4
Apple M1 Max
64.0 GB
400.0 GB/s
q6
36.84 GB
~7.6
~11.4 q4
Apple M2 Max
96.0 GB
400.0 GB/s
fp8
49.12 GB
~5.7
~11.4 q4
Apple M3 Max
128.0 GB
400.0 GB/s
fp8
49.12 GB
~5.7
~11.4 q4
Apple M4 Pro
64.0 GB
273.0 GB/s
q6
36.84 GB
~5.2
~7.8 q4
Where it runs
Taken from the card's availability section.
Platform
Id on that platform
Hugging Face
moonshotai/Kimi-Linear-48B-A3B-Instruct
Verified benchmark evidence
Each row was checked against its source by a reviewer, and carries the model identifier as actually evaluated — which is not always the same as this card's.