Granite 4.2 3B

IBM · ibm/granite-4-2-3b

ProviderIBM
Familygranite
Typellm-chat
Statusactive
Released2026-08-07
Updated2026-09-02
Parameters3.7B
Open weightsyes

Lineage

RelationshipModel
Is a finetune ofibm-granite/granite-4.1-3b-base

What it fits on

Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.

Every figure here is computed, not measured. Fit is weights at each quantisation against device memory, with a 25% allowance for the KV cache, activations and the OS. Decode rate is the memory-bandwidth roofline at 70% efficiency. Nobody has run this model on these devices.
DeviceMemoryBandwidthBest qualityWeightstok/sFastest
Cerebras WSE-344.0 GB21000000.0 GB/sbf167.32 GB~2008340.7~8033362.8 q4
NVIDIA GB200 Grace Blackwell Superchip372.0 GB16000.0 GB/sbf167.32 GB~1530.2~6120.7 q4
AMD Instinct MI355X288.0 GB8000.0 GB/sbf167.32 GB~765.1~3060.3 q4
NVIDIA B200180.0 GB7700.0 GB/sbf167.32 GB~736.4~2945.6 q4
Google TPU7x (Ironwood)192.0 GB7380.0 GB/sbf167.32 GB~705.8~2823.2 q4
AMD Instinct MI325X256.0 GB6000.0 GB/sbf167.32 GB~573.8~2295.2 q4
AMD Instinct MI300X192.0 GB5300.0 GB/sbf167.32 GB~506.9~2027.5 q4
NVIDIA H200 SXM141.0 GB4800.0 GB/sbf167.32 GB~459.0~1836.2 q4
NVIDIA H100 NVL94.0 GB3900.0 GB/sbf167.32 GB~373.0~1491.9 q4
NVIDIA H100 SXM80.0 GB3350.0 GB/sbf167.32 GB~320.4~1281.5 q4
AMD Instinct MI250X128.0 GB3200.0 GB/sbf167.32 GB~306.0~1224.1 q4
Google TPU v5p95.0 GB2765.0 GB/sbf167.32 GB~264.4~1057.7 q4
NVIDIA A100 80GB SXM80.0 GB2039.0 GB/sbf167.32 GB~195.0~780.0 q4
NVIDIA H100 PCIe80.0 GB2000.0 GB/sbf167.32 GB~191.3~765.1 q4
NVIDIA GeForce RTX 509032.0 GB1792.0 GB/sbf167.32 GB~171.4~685.5 q4
Google TPU v6e (Trillium)32.0 GB1638.0 GB/sbf167.32 GB~156.7~626.6 q4
AMD Instinct MI21064.0 GB1600.0 GB/sbf167.32 GB~153.0~612.1 q4
NVIDIA A100 40GB SXM40.0 GB1555.0 GB/sbf167.32 GB~148.7~594.9 q4
Google TPU v432.0 GB1200.0 GB/sbf167.32 GB~114.8~459.0 q4
NVIDIA GeForce RTX 409024.0 GB1008.0 GB/sbf167.32 GB~96.4~385.6 q4
AMD Radeon RX 7900 XTX24.0 GB960.0 GB/sbf167.32 GB~91.8~367.2 q4
NVIDIA GeForce RTX 508016.0 GB960.0 GB/sbf167.32 GB~91.8~367.2 q4
NVIDIA GeForce RTX 309024.0 GB936.0 GB/sbf167.32 GB~89.5~358.1 q4
NVIDIA L40S48.0 GB864.0 GB/sbf167.32 GB~82.6~330.5 q4
Apple M3 Ultra512.0 GB819.0 GB/sbf167.32 GB~78.3~313.3 q4
AMD Radeon RX 7900 XT20.0 GB800.0 GB/sbf167.32 GB~76.5~306.0 q4
Apple M2 Ultra192.0 GB800.0 GB/sbf167.32 GB~76.5~306.0 q4
Google TPU v5e16.0 GB800.0 GB/sbf167.32 GB~76.5~306.0 q4
NVIDIA GeForce RTX 4080 SUPER16.0 GB736.0 GB/sbf167.32 GB~70.4~281.6 q4
NVIDIA GeForce RTX 4070 Ti SUPER16.0 GB672.0 GB/sbf167.32 GB~64.3~257.1 q4
AMD Radeon RX 9070 XT16.0 GB640.0 GB/sbf167.32 GB~61.2~244.8 q4
Apple M4 Max128.0 GB546.0 GB/sbf167.32 GB~52.2~208.9 q4
Apple M1 Max64.0 GB400.0 GB/sbf167.32 GB~38.3~153.0 q4
Apple M2 Max96.0 GB400.0 GB/sbf167.32 GB~38.3~153.0 q4
Apple M3 Max128.0 GB400.0 GB/sbf167.32 GB~38.3~153.0 q4
NVIDIA GeForce RTX 3060 12GB12.0 GB360.0 GB/sbf167.32 GB~34.4~137.7 q4
NVIDIA L424.0 GB300.0 GB/sbf167.32 GB~28.7~114.8 q4
NVIDIA GeForce RTX 4060 Ti 16GB16.0 GB288.0 GB/sbf167.32 GB~27.5~110.2 q4
Apple M4 Pro64.0 GB273.0 GB/sbf167.32 GB~26.1~104.4 q4
Apple M432.0 GB120.0 GB/sbf167.32 GB~11.5~45.9 q4

Where it runs

Taken from the card's availability section.

PlatformId on that platform
Hugging Faceibm-granite/granite-4.2-3b

Verified benchmark evidence

Each row was checked against its source by a reviewer, and carries the model identifier as actually evaluated — which is not always the same as this card's.
BenchmarkEvaluated asScoreEvidence dateSource kind
aa_lcrGranite 4.2 3B24.33%2026-09-10 evaluatedindependent evaluatorsource
critptGranite 4.2 3B0.0%2026-09-10 evaluatedindependent evaluatorsource
gdpval_aaGranite 4.2 3B0.0%2026-09-10 evaluatedindependent evaluatorsource
gpqa_diamondGranite 4.2 3B55.86%2026-09-10 evaluatedindependent evaluatorsource
scicodeGranite 4.2 3B25.35%2026-09-10 evaluatedindependent evaluatorsource

Reported benchmark scores

This card reports no benchmark scores yet.

BenchmarkCatalogue standingScore

Data

This card as JSON · See it in the graph · Edit on GitHub