Cogito v2.1 671B

Deep Cogito · deepcogito/cogito-671b-v2-1

ProviderDeep Cogito
Familycogito
Typellm-reasoning
Statusactive
Released2025-11-19
Updated2025-11-21
Parameters671B
Open weightsyes

Lineage

RelationshipModel
Is a finetune ofdeepseek-ai/DeepSeek-V3-Base

Capabilities

Chain Of Thought · tier-2 Function Calling · tier-2 Multilingual · tier-2 Parallel Tool Calls · tier-2 Think Budget Control · tier-2

What it fits on

Ordered by predicted speed. Bandwidth sets decode rate; memory decides whether it runs at all.

Every figure here is computed, not measured. Fit is weights at each quantisation against device memory, with a 25% allowance for the KV cache, activations and the OS. Decode rate is the memory-bandwidth roofline at 70% efficiency. Nobody has run this model on these devices. This is a mixture-of-experts model and no card carries active-parameter counts, so these predictions use total parameters and understate the real speed.
DeviceMemoryBandwidthBest qualityWeightstok/sFastest
Apple M3 Ultra512.0 GB819.0 GB/sq4335.51 GB~1.7~1.7 q4

Where it runs

Taken from the card's availability section.

PlatformId on that platform
Hugging Facedeepcogito/cogito-671b-v2.1
Together AIdeepcogito/cogito-v2-1-671b

Verified benchmark evidence

Each row was checked against its source by a reviewer, and carries the model identifier as actually evaluated — which is not always the same as this card's.
BenchmarkEvaluated asScoreEvidence dateSource kind
aa_lcrCogito v2.122.67%2026-09-10 evaluatedindependent evaluatorsource
critptCogito v2.10.0%2026-09-10 evaluatedindependent evaluatorsource
gpqa_diamondCogito v2.176.77%2026-09-10 evaluatedindependent evaluatorsource

Reported benchmark scores

This card reports no benchmark scores yet.

BenchmarkCatalogue standingScore

Data

This card as JSON · See it in the graph · Edit on GitHub