Ternary Bonsai 27B

Prism ML · prism-ml/ternary-bonsai-27b

ProviderPrism ML
Familybonsai
Typequantized-variant
Statusactive
Released2026-07-14
Updated2026-08-31
Parameters27.4B
Open weightsyes

Lineage

RelationshipModel
Is a quantized ofQwen3.6 27B

Capabilities

Chain Of Thought · tier-2 Function Calling · tier-2

What it fits on

Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.

Every figure here is computed, not measured. Fit is weights at each quantisation against device memory, with a 25% allowance for the KV cache, activations and the OS. Decode rate is the memory-bandwidth roofline at 70% efficiency. Nobody has run this model on these devices.
DeviceMemoryBandwidthBest qualityWeightstok/sFastest
Cerebras WSE-344.0 GB21000000.0 GB/sfp827.36 GB~537345.0~1074689.9 q4
NVIDIA GB200 Grace Blackwell Superchip372.0 GB16000.0 GB/sbf1654.71 GB~204.7~818.8 q4
AMD Instinct MI355X288.0 GB8000.0 GB/sbf1654.71 GB~102.4~409.4 q4
NVIDIA B200180.0 GB7700.0 GB/sbf1654.71 GB~98.5~394.1 q4
Google TPU7x (Ironwood)192.0 GB7380.0 GB/sbf1654.71 GB~94.4~377.7 q4
AMD Instinct MI325X256.0 GB6000.0 GB/sbf1654.71 GB~76.8~307.1 q4
AMD Instinct MI300X192.0 GB5300.0 GB/sbf1654.71 GB~67.8~271.2 q4
NVIDIA H200 SXM141.0 GB4800.0 GB/sbf1654.71 GB~61.4~245.6 q4
NVIDIA H100 NVL94.0 GB3900.0 GB/sbf1654.71 GB~49.9~199.6 q4
NVIDIA H100 SXM80.0 GB3350.0 GB/sbf1654.71 GB~42.9~171.4 q4
AMD Instinct MI250X128.0 GB3200.0 GB/sbf1654.71 GB~40.9~163.8 q4
Google TPU v5p95.0 GB2765.0 GB/sbf1654.71 GB~35.4~141.5 q4
NVIDIA A100 80GB SXM80.0 GB2039.0 GB/sbf1654.71 GB~26.1~104.3 q4
NVIDIA H100 PCIe80.0 GB2000.0 GB/sbf1654.71 GB~25.6~102.4 q4
NVIDIA GeForce RTX 509032.0 GB1792.0 GB/sq620.52 GB~61.1~91.7 q4
Google TPU v6e (Trillium)32.0 GB1638.0 GB/sq620.52 GB~55.9~83.8 q4
AMD Instinct MI21064.0 GB1600.0 GB/sfp827.36 GB~40.9~81.9 q4
NVIDIA A100 40GB SXM40.0 GB1555.0 GB/sfp827.36 GB~39.8~79.6 q4
Google TPU v432.0 GB1200.0 GB/sq620.52 GB~40.9~61.4 q4
NVIDIA GeForce RTX 409024.0 GB1008.0 GB/sq517.1 GB~41.3~51.6 q4
AMD Radeon RX 7900 XTX24.0 GB960.0 GB/sq517.1 GB~39.3~49.1 q4
NVIDIA GeForce RTX 309024.0 GB936.0 GB/sq517.1 GB~38.3~47.9 q4
NVIDIA L40S48.0 GB864.0 GB/sfp827.36 GB~22.1~44.2 q4
Apple M3 Ultra512.0 GB819.0 GB/sbf1654.71 GB~10.5~41.9 q4
AMD Radeon RX 7900 XT20.0 GB800.0 GB/sq413.68 GB~40.9~40.9 q4
Apple M2 Ultra192.0 GB800.0 GB/sbf1654.71 GB~10.2~40.9 q4
Apple M4 Max128.0 GB546.0 GB/sbf1654.71 GB~7.0~27.9 q4
Apple M1 Max64.0 GB400.0 GB/sfp827.36 GB~10.2~20.5 q4
Apple M2 Max96.0 GB400.0 GB/sbf1654.71 GB~5.1~20.5 q4
Apple M3 Max128.0 GB400.0 GB/sbf1654.71 GB~5.1~20.5 q4
NVIDIA L424.0 GB300.0 GB/sq517.1 GB~12.3~15.4 q4
Apple M4 Pro64.0 GB273.0 GB/sfp827.36 GB~7.0~14.0 q4
Apple M432.0 GB120.0 GB/sq620.52 GB~4.1~6.1 q4

Where it runs

Taken from the card's availability section.

PlatformId on that platform
Hugging Faceprism-ml/Ternary-Bonsai-27B-gguf

Reported benchmark scores

This card reports no benchmark scores yet.

BenchmarkCatalogue standingScore

Data

This card as JSON · See it in the graph · Edit on GitHub