Claude Mythos Preview

Anthropic · anthropic/claude-mythos-preview

ProviderAnthropic
Familyclaude-mythos
Typellm-reasoning
Statuspreview
Released2026-04-07
Updated2026-04-07

Reported benchmark scores

These scores carry one collection date for the whole card (2026-04) and a source list rather than a source per score, so they cannot be attributed individually. They are shown as reported and are not verified evidence. Benchmark pages carry the reviewed, dated evidence where it exists.
BenchmarkCatalogue standingScore
Arena Elo — Overall (Text)unassessed1410.0
BrowseCompunassessed86.9
CharXiv Reasoningunassessed86.1
CharXiv Reasoning (with tool use)unassessed93.2
GPQA Diamondunassessed94.55
GraphWalks BFS (256K-1M context)unassessed80.0
GraphWalks Parents (256K-1M context)unassessed97.7
Humanity's Last Examunassessed56.8
Humanity's Last Exam (with tools)unassessed64.7
HumanEvalunassessed93.2
IFEvalunassessed92.1
LAB-Bench: FigQAunassessed79.7
LAB-Bench: FigQA (with tools)unassessed89.0
MATH-500unassessed96.4
MMLU-Prounassessed85.2
Multilingual MMLU (MMMLU)unassessed92.67
MT-Benchunassessed9.4
OSWorldunassessed79.6
ScreenSpot-Prounassessed79.5
ScreenSpot-Pro (with tools)unassessed92.8
SWE-bench Multilingualunassessed87.3
SWE-bench Multimodalunassessed59.0
SWE-bench Prounassessed77.8
SWE-bench Verifiedunassessed93.9
Terminal-Bench 2.0unassessed82.0
USAMO 2026unassessed97.6

Data

This card as JSON · Edit on GitHub