Llama 3.1 Nemotron 70B

NVIDIA · nvidia/llama-3-1-nemotron-70b-instruct-hf

ProviderNVIDIA
Familyllama-nemotron
Typellm-chat
Statusactive
Released2024-10-12
Updated2025-04-13
Parameters70.6B
Open weightsyes

Lineage

RelationshipModel
Is a finetune ofmeta-llama/Llama-3.1-70B-Instruct

What it fits on

Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.

Every figure here is computed, not measured. Fit is weights at each quantisation against device memory, with a 25% allowance for the KV cache, activations and the OS. Decode rate is the memory-bandwidth roofline at 70% efficiency. Nobody has run this model on these devices.
DeviceMemoryBandwidthBest qualityWeightstok/sFastest
NVIDIA GB200 Grace Blackwell Superchip372.0 GB16000.0 GB/sbf16141.11 GB~79.4~317.5 q4
AMD Instinct MI355X288.0 GB8000.0 GB/sbf16141.11 GB~39.7~158.7 q4
NVIDIA B200180.0 GB7700.0 GB/sfp870.55 GB~76.4~152.8 q4
Google TPU7x (Ironwood)192.0 GB7380.0 GB/sbf16141.11 GB~36.6~146.4 q4
AMD Instinct MI325X256.0 GB6000.0 GB/sbf16141.11 GB~29.8~119.1 q4
AMD Instinct MI300X192.0 GB5300.0 GB/sbf16141.11 GB~26.3~105.2 q4
NVIDIA H200 SXM141.0 GB4800.0 GB/sfp870.55 GB~47.6~95.2 q4
NVIDIA H100 NVL94.0 GB3900.0 GB/sq652.92 GB~51.6~77.4 q4
NVIDIA H100 SXM80.0 GB3350.0 GB/sq652.92 GB~44.3~66.5 q4
AMD Instinct MI250X128.0 GB3200.0 GB/sfp870.55 GB~31.7~63.5 q4
Google TPU v5p95.0 GB2765.0 GB/sfp870.55 GB~27.4~54.9 q4
NVIDIA A100 80GB SXM80.0 GB2039.0 GB/sq652.92 GB~27.0~40.5 q4
NVIDIA H100 PCIe80.0 GB2000.0 GB/sq652.92 GB~26.5~39.7 q4
AMD Instinct MI21064.0 GB1600.0 GB/sq544.1 GB~25.4~31.7 q4
NVIDIA L40S48.0 GB864.0 GB/sq435.28 GB~17.1~17.1 q4
Apple M3 Ultra512.0 GB819.0 GB/sbf16141.11 GB~4.1~16.3 q4
Apple M2 Ultra192.0 GB800.0 GB/sbf16141.11 GB~4.0~15.9 q4
Apple M4 Max128.0 GB546.0 GB/sfp870.55 GB~5.4~10.8 q4
Apple M1 Max64.0 GB400.0 GB/sq544.1 GB~6.3~7.9 q4
Apple M2 Max96.0 GB400.0 GB/sfp870.55 GB~4.0~7.9 q4
Apple M3 Max128.0 GB400.0 GB/sfp870.55 GB~4.0~7.9 q4
Apple M4 Pro64.0 GB273.0 GB/sq544.1 GB~4.3~5.4 q4

Where it runs

Taken from the card's availability section.

PlatformId on that platform
Hugging Facenvidia/Llama-3.1-Nemotron-70B-Instruct-HF

Verified benchmark evidence

Each row was checked against its source by a reviewer, and carries the model identifier as actually evaluated — which is not always the same as this card's.
BenchmarkEvaluated asScoreEvidence dateSource kind
aa_lcrLlama 3.1 Nemotron 70B8.33%2026-09-10 evaluatedindependent evaluatorsource
critptLlama 3.1 Nemotron 70B0.0%2026-09-10 evaluatedindependent evaluatorsource
gpqa_diamondLlama 3.1 Nemotron 70B46.46%2026-09-10 evaluatedindependent evaluatorsource

Reported benchmark scores

This card reports no benchmark scores yet.

BenchmarkCatalogue standingScore

Data

This card as JSON · See it in the graph · Edit on GitHub