ModelSpec

How ModelSpec decides.

The front page makes five claims. This page takes each one apart in plain language, and links it to the code, the document or the public data that makes it true. Check it rather than trust it.

Everything here describes the code on main. Where a figure is needed, the page shows how it is computed, or reads it from the live snapshot.

From a published score to your answer

  1. 1SourcesModel cards and benchmark pages. Each value has its URL, its date and who measured it.
  2. 2Two keysA second, independent check must agree before a value is used.
  3. 3SnapshotEverything admitted, frozen into one file and hashed.
  4. 4EstimateAbility per domain, learned from every benchmark, with an 80% range.
  5. 5Your specMust gates remove. Prefer weights order. Unknowns stay listed.
  6. 6AnswerRanked where the evidence separates models, and flagged where it can't.
1 · Where evidence comes from

Every number starts as a sourced, checked record.

A value with no source never reaches a decision. Neither does one that only its collector has checked.

Cards and benchmark pages

Model cards hold the measurements, one record per result. Benchmark pages say what each benchmark measures, and which domains it counts toward.

Who measured it

Every measurement carries one of three labels. A lab's claim about its own model sits beside independent results, labelled, never mixed in unmarked.

benchmark_authorThe benchmark's own authors.independent_evaluatorAn evaluator outside the lab.provider_self_reportThe lab or provider, about its own model.

Two keys before anything counts

A record enters the snapshot only when a second check agrees with the first, and that check was made by a deterministic tool or by a model from a different family.

What never enters

A value whose latest check isn't "verified", anything without a registered source URL, anything from an excluded source, or the old flat score block on each card.

Sources we exclude

Two publishers whose terms don't permit our use were removed on 24 September 2026, with every value and benchmark they own. A test fails the build if either comes back, and the snapshot build scans its own output with the same rules.

Check it

What evidence_basis means

The rank export and the /v1/rank API label each row's inputs with evidence_basis. The label describes the inputs. It is not a verdict on the model.

noneNo usable measurement contributed.unverified-legacyEvery input is an older card value.mixedReviewed and older values both contribute.partial-verifiedEvery input is reviewed, but some the profile asks for are missing.verifiedEvery weighted benchmark is present and reviewed.

Check it

2 · The capability estimate

Ability is estimated per domain, from every benchmark we hold.

There is no fixed benchmark list and no hand-set weight. Every admitted benchmark with at least two model observations counts, and nobody picks favourites. One leaderboard is one reading, and readings disagree. The estimate uses all of them, and says how sure it is.

No benchmark names in the code

The fit reads which domains a benchmark counts toward from that benchmark's own page, tagged direct or proxy.

Direct counts more than proxy

A direct measurement loads at 1, a proxy at 0.35, so a proxy carries 0.1225 of a direct one's precision.

Old evidence counts less

A measurement's precision halves every 365 days, down to a floor of a quarter.

Every estimate has a range

A central 80% range: the estimate ± 1.2816 standard deviations.

No evidence, no estimate

A model with nothing measured in a domain gets no estimate there.

A higher score never hurts

The weights don't depend on the scores.

Check it

3 · Ties

Why the #1 is often a tie.

Every estimate is a range. When two models' ranges overlap, the evidence can't say which is better, so the page doesn't pretend to.

range         = estimate ± 1.2816 × sd
not_separable = max(A.low, B.low) ≤ min(A.high, B.high)
                for any other model B
the top estimate's lower boundClaude Opus 5.5leaderClaude Sonnet 5.5tiedClaude Opus 4.7tiedClaude Fable 5tiedGPT-6 AstratiedClaude Fable 5.1tiedClaude Opus 5tiedGPT-6 Soltiedestimated ability in one domain, higher is better
18 of the other 21 models reach Claude Opus 5.5's lower bound. That is the front-page figure, counted against the top estimate only. Of those, Gemini 3.7 Flash is cheapest.

Overlap with any other model

A row is flagged not_separable when its range overlaps the range of any other model in the domain being ranked, not only the leader's. Two models can be told apart from the leader and still not from each other. The board names the models the evidence cannot separate. The API keeps a stable transport order so a list can travel. Read the flag before the rank.

Inside a tie, choose on something else

Ability cannot split a tied group. Price, context, licence and the other facets you set can, and the board shows them for every row.

Check it

4 · Must and Prefer

Must is a gate. Prefer is a weight.

Every facet on the board has three settings, and each does a different job. Conditions filter. They never add points.

Doesn't matter

The default. The row stays on the board, greyed, so you can always see what you didn't choose.

Must

A hard gate, with a threshold where the facet has one: context of at least 200K, a cost cap, a signed BAA. A model that fails is excluded, and stays on screen as excluded.

Prefer

A weight. It orders the models that passed every Must. It never removes one.

For agents, in a spec. Musts are the conditions under where. Prefers are the weights under optimize. A condition marked soft(penalty) is a preference too.

Check it

5 · Unknown means unknown

A missing fact is never a zero.

Every condition has three answers: pass, fail and unknown. Unknown is its own answer, with its own rules. A null beats a guess.

The value isWhich meansWhat the answer does
KnownA value with a source, and a check that passed.Used.
Unknown, on ability or a spec such as contextNobody has published it, or it hasn't been collected yet.May qualify: listed beside the answer, not ranked, never scored as zero.
Unknown, on your data or the licenceWe can't confirm the provider's terms.Not treated as met. Shown as unverified, may qualify, so you can check it yourself.
InapplicableThe model's class can't have it. A model that writes no text has no output-token limit.Derived from the class, never typed on a card.

Check it

6 · Reproducibility

Same spec, same snapshot, same answer.

A decision depends on two things you can name: the question and the evidence. Pin both, and anyone holding the same snapshot file gets the answer you got.

The snapshot is fixed

Every admitted record, written as canonical JSON: sorted keys, no whitespace.

The spec is hashed too

Key order, whitespace and compact or YAML spelling don't change the hash. Changing any field does.

The decision id comes from both

dec_ plus a hash of the spec hash and the snapshot id.

Signing state at build

This page reads the signature block from the decision snapshot that the site publishes.

reproduce a decision, offline
$ modelspec snapshot fetch
$ modelspec decide spec.yaml --json
  "snapshot": "[SNAPSHOT ID]",
  "spec_hash": "sha256:[SPEC HASH]",
  "signature_verified": true

Public signing key

This snapshot is Ed25519-signed with published key ID ed25519-dd3ae9b0899184e5. The signature was verified when this page was built, with the same check the CLI runs.

Run modelspec snapshot fetch to download the snapshot and verify its signature yourself.

The content hash is checked separately.

Check it

7 · Neutrality

Nobody pays to rank higher.

We charge the people and agents who ask for an answer. Never the models, labs or hosts the answer is about.

“No referral fees, no paid placement, no provider-paid visibility, permanently.”
The neutrality commitment, neutrality-v1
  • accepts_referral_feesfalse
  • accepts_paid_placementfalse
  • accepts_provider_paid_visibilityfalse
  • proxies_inference_tokensfalse
  • stores_customer_promptsfalse
  • conceals_purchases_from_catalogued_vendorsfalse
  • lets_supplier_models_write_supplier_cardsfalse

Check it

8 · What we don't do

What we don't do, on purpose.

No text box on the board

You set each criterion yourself. No parser and no AI reads your words and guesses what you meant.

No paid placement or referral fees

No provider, host or gateway can buy inclusion, position or a mention.

No proxying your tokens

We recommend and hand off. Your inference goes to the provider directly.

No prompts kept

A request carries a profile, not prompt text, so there is nothing to keep.

No guessed values

A missing value stays missing. It is never filled in, and never counted as zero.

No fixed benchmark list

Nobody chooses which benchmarks matter. Every admitted benchmark counts, weighted by the same rules.

Check it yourself.

The decision snapshot and vocabulary read by decision answers are public, versioned and free to fetch. So is the code that reads them. The /v1/policy-check determinations are private.

Open the board