DiffusionGemma 26B A4B

Google · google/diffusiongemma-26b-a4b-it

ProviderGoogle
Familygemma
Typeimage-generation
Statusactive
Released2026-06-09
Updated2026-07-15
Parameters25.8B
Open weightsyes

What it fits on

Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.

Every figure here is computed, not measured. Fit is weights at each quantisation against device memory, with a 25% allowance for the KV cache, activations and the OS. Decode rate is the memory-bandwidth roofline at 70% efficiency. Nobody has run this model on these devices. This is a mixture-of-experts model and no card carries active-parameter counts, so these predictions use total parameters and understate the real speed.
DeviceMemoryBandwidthBest qualityWeightstok/sFastest
Cerebras WSE-344.0 GB21000000.0 GB/sfp825.82 GB~569242.8~1138485.6 q4
NVIDIA GB200 Grace Blackwell Superchip372.0 GB16000.0 GB/sbf1651.65 GB~216.9~867.4 q4
AMD Instinct MI355X288.0 GB8000.0 GB/sbf1651.65 GB~108.4~433.7 q4
NVIDIA B200180.0 GB7700.0 GB/sbf1651.65 GB~104.4~417.4 q4
Google TPU7x (Ironwood)192.0 GB7380.0 GB/sbf1651.65 GB~100.0~400.1 q4
AMD Instinct MI325X256.0 GB6000.0 GB/sbf1651.65 GB~81.3~325.3 q4
AMD Instinct MI300X192.0 GB5300.0 GB/sbf1651.65 GB~71.8~287.3 q4
NVIDIA H200 SXM141.0 GB4800.0 GB/sbf1651.65 GB~65.1~260.2 q4
NVIDIA H100 NVL94.0 GB3900.0 GB/sbf1651.65 GB~52.9~211.4 q4
NVIDIA H100 SXM80.0 GB3350.0 GB/sbf1651.65 GB~45.4~181.6 q4
AMD Instinct MI250X128.0 GB3200.0 GB/sbf1651.65 GB~43.4~173.5 q4
Google TPU v5p95.0 GB2765.0 GB/sbf1651.65 GB~37.5~149.9 q4
NVIDIA A100 80GB SXM80.0 GB2039.0 GB/sbf1651.65 GB~27.6~110.5 q4
NVIDIA H100 PCIe80.0 GB2000.0 GB/sbf1651.65 GB~27.1~108.4 q4
NVIDIA GeForce RTX 509032.0 GB1792.0 GB/sq619.37 GB~64.8~97.2 q4
Google TPU v6e (Trillium)32.0 GB1638.0 GB/sq619.37 GB~59.2~88.8 q4
AMD Instinct MI21064.0 GB1600.0 GB/sfp825.82 GB~43.4~86.7 q4
NVIDIA A100 40GB SXM40.0 GB1555.0 GB/sfp825.82 GB~42.2~84.3 q4
Google TPU v432.0 GB1200.0 GB/sq619.37 GB~43.4~65.1 q4
NVIDIA GeForce RTX 409024.0 GB1008.0 GB/sq516.14 GB~43.7~54.6 q4
AMD Radeon RX 7900 XTX24.0 GB960.0 GB/sq516.14 GB~41.6~52.0 q4
NVIDIA GeForce RTX 309024.0 GB936.0 GB/sq516.14 GB~40.6~50.7 q4
NVIDIA L40S48.0 GB864.0 GB/sfp825.82 GB~23.4~46.8 q4
Apple M3 Ultra512.0 GB819.0 GB/sbf1651.65 GB~11.1~44.4 q4
AMD Radeon RX 7900 XT20.0 GB800.0 GB/sq412.91 GB~43.4~43.4 q4
Apple M2 Ultra192.0 GB800.0 GB/sbf1651.65 GB~10.8~43.4 q4
Apple M4 Max128.0 GB546.0 GB/sbf1651.65 GB~7.4~29.6 q4
Apple M1 Max64.0 GB400.0 GB/sfp825.82 GB~10.8~21.7 q4
Apple M2 Max96.0 GB400.0 GB/sbf1651.65 GB~5.4~21.7 q4
Apple M3 Max128.0 GB400.0 GB/sbf1651.65 GB~5.4~21.7 q4
NVIDIA L424.0 GB300.0 GB/sq516.14 GB~13.0~16.3 q4
Apple M4 Pro64.0 GB273.0 GB/sfp825.82 GB~7.4~14.8 q4
Apple M432.0 GB120.0 GB/sq619.37 GB~4.3~6.5 q4

Where it runs

Taken from the card's availability section.

PlatformId on that platform
Hugging Facegoogle/diffusiongemma-26B-A4B-it

Verified benchmark evidence

Each row was checked against its source by a reviewer, and carries the model identifier as actually evaluated — which is not always the same as this card's.
BenchmarkEvaluated asScoreEvidence dateSource kind
aa_lcrDiffusionGemma 26B A4B19.67%2026-09-10 evaluatedindependent evaluatorsource
critptDiffusionGemma 26B A4B0.29%2026-09-10 evaluatedindependent evaluatorsource
gdpval_aaDiffusionGemma 26B A4B0.06%2026-09-10 evaluatedindependent evaluatorsource
gpqa_diamondDiffusionGemma 26B A4B66.87%2026-09-10 evaluatedindependent evaluatorsource

Reported benchmark scores

This card reports no benchmark scores yet.

BenchmarkCatalogue standingScore

Data

This card as JSON · See it in the graph · Edit on GitHub