Alibaba Cloud · qwen/qwen2-72b-instruct
| Provider | Alibaba Cloud |
|---|---|
| Family | qwen |
| Type | llm-chat |
| Status | active |
| Released | 2024-05-28 |
| Updated | 2024-10-08 |
| Parameters | 72.7B |
| Open weights | yes |
| Relationship | Model |
|---|---|
| Is a finetune of | Qwen/Qwen2-72B |
Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.
| Device | Memory | Bandwidth | Best quality | Weights | tok/s | Fastest |
|---|---|---|---|---|---|---|
| NVIDIA GB200 Grace Blackwell Superchip | 372.0 GB | 16000.0 GB/s | bf16 | 145.41 GB | ~77.0 | ~308.1 q4 |
| AMD Instinct MI355X | 288.0 GB | 8000.0 GB/s | bf16 | 145.41 GB | ~38.5 | ~154.0 q4 |
| NVIDIA B200 | 180.0 GB | 7700.0 GB/s | fp8 | 72.71 GB | ~74.1 | ~148.3 q4 |
| Google TPU7x (Ironwood) | 192.0 GB | 7380.0 GB/s | fp8 | 72.71 GB | ~71.1 | ~142.1 q4 |
| AMD Instinct MI325X | 256.0 GB | 6000.0 GB/s | bf16 | 145.41 GB | ~28.9 | ~115.5 q4 |
| AMD Instinct MI300X | 192.0 GB | 5300.0 GB/s | fp8 | 72.71 GB | ~51.0 | ~102.1 q4 |
| NVIDIA H200 SXM | 141.0 GB | 4800.0 GB/s | fp8 | 72.71 GB | ~46.2 | ~92.4 q4 |
| NVIDIA H100 NVL | 94.0 GB | 3900.0 GB/s | q6 | 54.53 GB | ~50.1 | ~75.1 q4 |
| NVIDIA H100 SXM | 80.0 GB | 3350.0 GB/s | q6 | 54.53 GB | ~43.0 | ~64.5 q4 |
| AMD Instinct MI250X | 128.0 GB | 3200.0 GB/s | fp8 | 72.71 GB | ~30.8 | ~61.6 q4 |
| Google TPU v5p | 95.0 GB | 2765.0 GB/s | q6 | 54.53 GB | ~35.5 | ~53.2 q4 |
| NVIDIA A100 80GB SXM | 80.0 GB | 2039.0 GB/s | q6 | 54.53 GB | ~26.2 | ~39.3 q4 |
| NVIDIA H100 PCIe | 80.0 GB | 2000.0 GB/s | q6 | 54.53 GB | ~25.7 | ~38.5 q4 |
| AMD Instinct MI210 | 64.0 GB | 1600.0 GB/s | q5 | 45.44 GB | ~24.6 | ~30.8 q4 |
| Apple M3 Ultra | 512.0 GB | 819.0 GB/s | bf16 | 145.41 GB | ~3.9 | ~15.8 q4 |
| Apple M2 Ultra | 192.0 GB | 800.0 GB/s | fp8 | 72.71 GB | ~7.7 | ~15.4 q4 |
| Apple M4 Max | 128.0 GB | 546.0 GB/s | fp8 | 72.71 GB | ~5.3 | ~10.5 q4 |
| Apple M1 Max | 64.0 GB | 400.0 GB/s | q5 | 45.44 GB | ~6.2 | ~7.7 q4 |
| Apple M2 Max | 96.0 GB | 400.0 GB/s | q6 | 54.53 GB | ~5.1 | ~7.7 q4 |
| Apple M3 Max | 128.0 GB | 400.0 GB/s | fp8 | 72.71 GB | ~3.9 | ~7.7 q4 |
| Apple M4 Pro | 64.0 GB | 273.0 GB/s | q5 | 45.44 GB | ~4.2 | ~5.3 q4 |
Taken from the card's availability section.
| Platform | Id on that platform |
|---|---|
| Hugging Face | Qwen/Qwen2-72B-Instruct |
This card reports no benchmark scores yet.
| Benchmark | Catalogue standing | Score |
|---|