NVIDIA Nemotron 3.5 Lightning 30B A3B

NVIDIA · nvidia/nvidia-nemotron-3-5-lightning-30b-a3b

ProviderNVIDIA
Familynemotron
Typellm-reasoning
Statusactive
Released2026-08-11
Updated2026-08-24
Parameters31.6B
Open weightsyes

Capabilities

Chain Of Thought · tier-2 Function Calling · tier-2 Multilingual · tier-2 Think Budget Control · tier-2

What it fits on

Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.

Every figure here is computed, not measured. Fit is weights at each quantisation against device memory, with a 25% allowance for the KV cache, activations and the OS. Decode rate is the memory-bandwidth roofline at 70% efficiency. Nobody has run this model on these devices.
DeviceMemoryBandwidthBest qualityWeightstok/sFastest
Cerebras WSE-344.0 GB21000000.0 GB/sfp831.58 GB~4900000.0~9800000.0 q4
NVIDIA GB200 Grace Blackwell Superchip372.0 GB16000.0 GB/sbf1663.16 GB~1866.7~7466.7 q4
AMD Instinct MI355X288.0 GB8000.0 GB/sbf1663.16 GB~933.3~3733.3 q4
NVIDIA B200180.0 GB7700.0 GB/sbf1663.16 GB~898.3~3593.3 q4
Google TPU7x (Ironwood)192.0 GB7380.0 GB/sbf1663.16 GB~861.0~3444.0 q4
AMD Instinct MI325X256.0 GB6000.0 GB/sbf1663.16 GB~700.0~2800.0 q4
AMD Instinct MI300X192.0 GB5300.0 GB/sbf1663.16 GB~618.3~2473.3 q4
NVIDIA H200 SXM141.0 GB4800.0 GB/sbf1663.16 GB~560.0~2240.0 q4
NVIDIA H100 NVL94.0 GB3900.0 GB/sbf1663.16 GB~455.0~1820.0 q4
NVIDIA H100 SXM80.0 GB3350.0 GB/sfp831.58 GB~781.7~1563.3 q4
AMD Instinct MI250X128.0 GB3200.0 GB/sbf1663.16 GB~373.3~1493.3 q4
Google TPU v5p95.0 GB2765.0 GB/sbf1663.16 GB~322.6~1290.3 q4
NVIDIA A100 80GB SXM80.0 GB2039.0 GB/sfp831.58 GB~475.8~951.5 q4
NVIDIA H100 PCIe80.0 GB2000.0 GB/sfp831.58 GB~466.7~933.3 q4
NVIDIA GeForce RTX 509032.0 GB1792.0 GB/sq623.68 GB~557.5~836.3 q4
Google TPU v6e (Trillium)32.0 GB1638.0 GB/sq623.68 GB~509.6~764.4 q4
AMD Instinct MI21064.0 GB1600.0 GB/sfp831.58 GB~373.3~746.7 q4
NVIDIA A100 40GB SXM40.0 GB1555.0 GB/sq623.68 GB~483.8~725.7 q4
Google TPU v432.0 GB1200.0 GB/sq623.68 GB~373.3~560.0 q4
NVIDIA GeForce RTX 409024.0 GB1008.0 GB/sq415.79 GB~470.4~470.4 q4
AMD Radeon RX 7900 XTX24.0 GB960.0 GB/sq415.79 GB~448.0~448.0 q4
NVIDIA GeForce RTX 309024.0 GB936.0 GB/sq415.79 GB~436.8~436.8 q4
NVIDIA L40S48.0 GB864.0 GB/sfp831.58 GB~201.6~403.2 q4
Apple M3 Ultra512.0 GB819.0 GB/sbf1663.16 GB~95.5~382.2 q4
Apple M2 Ultra192.0 GB800.0 GB/sbf1663.16 GB~93.3~373.3 q4
Apple M4 Max128.0 GB546.0 GB/sbf1663.16 GB~63.7~254.8 q4
Apple M1 Max64.0 GB400.0 GB/sfp831.58 GB~93.3~186.7 q4
Apple M2 Max96.0 GB400.0 GB/sbf1663.16 GB~46.7~186.7 q4
Apple M3 Max128.0 GB400.0 GB/sbf1663.16 GB~46.7~186.7 q4
NVIDIA L424.0 GB300.0 GB/sq415.79 GB~140.0~140.0 q4
Apple M4 Pro64.0 GB273.0 GB/sfp831.58 GB~63.7~127.4 q4
Apple M432.0 GB120.0 GB/sq623.68 GB~37.3~56.0 q4

Where it runs

Taken from the card's availability section.

PlatformId on that platform
Fireworks AI
Hugging Facenvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
NVIDIA NIMnvidia/nemotron-3.5-lightning-30b-a3b
OpenRouternvidia/nemotron-3.5-lightning

Verified benchmark evidence

Each row was checked against its source by a reviewer, and carries the model identifier as actually evaluated — which is not always the same as this card's.
BenchmarkEvaluated asScoreEvidence dateSource kind
aa_lcrNemotron 3.5 Lightning60.33%2026-09-10 evaluatedindependent evaluatorsource
critptNemotron 3.5 Lightning0.0%2026-09-10 evaluatedindependent evaluatorsource
gdpval_aaNemotron 3.5 Lightning13.34%2026-09-10 evaluatedindependent evaluatorsource
gpqa_diamondNemotron 3.5 Lightning74.34%2026-09-10 evaluatedindependent evaluatorsource
scicodeNemotron 3.5 Lightning32.06%2026-09-10 evaluatedindependent evaluatorsource

Reported benchmark scores

This card reports no benchmark scores yet.

BenchmarkCatalogue standingScore

Data

This card as JSON · See it in the graph · Edit on GitHub