Qwen2 72B Instruct

Alibaba Cloud · qwen/qwen2-72b-instruct

ProviderAlibaba Cloud
Familyqwen
Typellm-chat
Statusactive
Released2024-05-28
Updated2024-10-08
Parameters72.7B
Open weightsyes

Lineage

RelationshipModel
Is a finetune ofQwen/Qwen2-72B

What it fits on

Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.

Every figure here is computed, not measured. Fit is weights at each quantisation against device memory, with a 25% allowance for the KV cache, activations and the OS. Decode rate is the memory-bandwidth roofline at 70% efficiency. Nobody has run this model on these devices.
DeviceMemoryBandwidthBest qualityWeightstok/sFastest
NVIDIA GB200 Grace Blackwell Superchip372.0 GB16000.0 GB/sbf16145.41 GB~77.0~308.1 q4
AMD Instinct MI355X288.0 GB8000.0 GB/sbf16145.41 GB~38.5~154.0 q4
NVIDIA B200180.0 GB7700.0 GB/sfp872.71 GB~74.1~148.3 q4
Google TPU7x (Ironwood)192.0 GB7380.0 GB/sfp872.71 GB~71.1~142.1 q4
AMD Instinct MI325X256.0 GB6000.0 GB/sbf16145.41 GB~28.9~115.5 q4
AMD Instinct MI300X192.0 GB5300.0 GB/sfp872.71 GB~51.0~102.1 q4
NVIDIA H200 SXM141.0 GB4800.0 GB/sfp872.71 GB~46.2~92.4 q4
NVIDIA H100 NVL94.0 GB3900.0 GB/sq654.53 GB~50.1~75.1 q4
NVIDIA H100 SXM80.0 GB3350.0 GB/sq654.53 GB~43.0~64.5 q4
AMD Instinct MI250X128.0 GB3200.0 GB/sfp872.71 GB~30.8~61.6 q4
Google TPU v5p95.0 GB2765.0 GB/sq654.53 GB~35.5~53.2 q4
NVIDIA A100 80GB SXM80.0 GB2039.0 GB/sq654.53 GB~26.2~39.3 q4
NVIDIA H100 PCIe80.0 GB2000.0 GB/sq654.53 GB~25.7~38.5 q4
AMD Instinct MI21064.0 GB1600.0 GB/sq545.44 GB~24.6~30.8 q4
Apple M3 Ultra512.0 GB819.0 GB/sbf16145.41 GB~3.9~15.8 q4
Apple M2 Ultra192.0 GB800.0 GB/sfp872.71 GB~7.7~15.4 q4
Apple M4 Max128.0 GB546.0 GB/sfp872.71 GB~5.3~10.5 q4
Apple M1 Max64.0 GB400.0 GB/sq545.44 GB~6.2~7.7 q4
Apple M2 Max96.0 GB400.0 GB/sq654.53 GB~5.1~7.7 q4
Apple M3 Max128.0 GB400.0 GB/sfp872.71 GB~3.9~7.7 q4
Apple M4 Pro64.0 GB273.0 GB/sq545.44 GB~4.2~5.3 q4

Where it runs

Taken from the card's availability section.

PlatformId on that platform
Hugging FaceQwen/Qwen2-72B-Instruct

Reported benchmark scores

This card reports no benchmark scores yet.

BenchmarkCatalogue standingScore

Data

This card as JSON · See it in the graph · Edit on GitHub