Llama Nemotron Ultra

NVIDIA · nvidia/llama-3-1-nemotron-ultra-253b-v1

ProviderNVIDIA
Familyllama-nemotron
Typellm-chat
Statusactive
Released2025-04-07
Updated2025-10-15
Parameters253.4B
Open weightsyes

What it fits on

Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.

Every figure here is computed, not measured. Fit is weights at each quantisation against device memory, with a 25% allowance for the KV cache, activations and the OS. Decode rate is the memory-bandwidth roofline at 70% efficiency. Nobody has run this model on these devices.
DeviceMemoryBandwidthBest qualityWeightstok/sFastest
NVIDIA GB200 Grace Blackwell Superchip372.0 GB16000.0 GB/sfp8253.4 GB~44.2~88.4 q4
AMD Instinct MI355X288.0 GB8000.0 GB/sq6190.05 GB~29.5~44.2 q4
NVIDIA B200180.0 GB7700.0 GB/sq4126.7 GB~42.5~42.5 q4
Google TPU7x (Ironwood)192.0 GB7380.0 GB/sq4126.7 GB~40.8~40.8 q4
AMD Instinct MI325X256.0 GB6000.0 GB/sq6190.05 GB~22.1~33.1 q4
AMD Instinct MI300X192.0 GB5300.0 GB/sq4126.7 GB~29.3~29.3 q4
Apple M3 Ultra512.0 GB819.0 GB/sfp8253.4 GB~2.3~4.5 q4
Apple M2 Ultra192.0 GB800.0 GB/sq4126.7 GB~4.4~4.4 q4

Where it runs

Taken from the card's availability section.

PlatformId on that platform
Hugging Facenvidia/Llama-3_1-Nemotron-Ultra-253B-v1

Verified benchmark evidence

Each row was checked against its source by a reviewer, and carries the model identifier as actually evaluated — which is not always the same as this card's.
BenchmarkEvaluated asScoreEvidence dateSource kind
critptLlama Nemotron Ultra0.0%2026-09-10 evaluatedindependent evaluatorsource
gpqa_diamondLlama Nemotron Ultra72.83%2026-09-10 evaluatedindependent evaluatorsource

Reported benchmark scores

This card reports no benchmark scores yet.

BenchmarkCatalogue standingScore

Data

This card as JSON · See it in the graph · Edit on GitHub