Hermes 4 405B

Nous Research · nous-research/hermes-4-405b

ProviderNous Research
Familyhermes
Typellm-chat
Statusactive
Released2025-08-06
Updated2025-09-02
Parameters405.9B
Open weightsyes

Lineage

RelationshipModel
Is a finetune ofmeta-llama/Llama-3.1-405B

What it fits on

Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.

Every figure here is computed, not measured. Fit is weights at each quantisation against device memory, with a 25% allowance for the KV cache, activations and the OS. Decode rate is the memory-bandwidth roofline at 70% efficiency. Nobody has run this model on these devices.
DeviceMemoryBandwidthBest qualityWeightstok/sFastest
NVIDIA GB200 Grace Blackwell Superchip372.0 GB16000.0 GB/sq5253.66 GB~44.2~55.2 q4
AMD Instinct MI355X288.0 GB8000.0 GB/sq4202.93 GB~27.6~27.6 q4
Apple M3 Ultra512.0 GB819.0 GB/sq6304.39 GB~1.9~2.8 q4

Where it runs

Taken from the card's availability section.

PlatformId on that platform
Hugging FaceNousResearch/Hermes-4-405B

Verified benchmark evidence

Each row was checked against its source by a reviewer, and carries the model identifier as actually evaluated — which is not always the same as this card's.
BenchmarkEvaluated asScoreEvidence dateSource kind
aa_lcrHermes 4 405B22.33%2026-09-10 evaluatedindependent evaluatorsource
critptHermes 4 405B0.29%2026-09-10 evaluatedindependent evaluatorsource
gpqa_diamondHermes 4 405B72.73%2026-09-10 evaluatedindependent evaluatorsource

Reported benchmark scores

This card reports no benchmark scores yet.

BenchmarkCatalogue standingScore

Data

This card as JSON · See it in the graph · Edit on GitHub