Two keys before anything counts
A record enters the snapshot only when a second check agrees with the first, and that check was made by a deterministic tool or by a model from a different family.
The front page makes five claims. This page takes each one apart in plain language, and links it to the code, the document or the public data that makes it true. Check it rather than trust it.
Everything here describes the code on main. Where a figure is needed, the page shows how it is computed, or reads it from the live snapshot.
1SourcesModel cards and benchmark pages. Each value has its URL, its date and who measured it.2Two keysA second, independent check must agree before a value is used.3SnapshotEverything admitted, frozen into one file and hashed.4EstimateAbility per domain, learned from every benchmark, with an 80% range.5Your specMust gates remove. Prefer weights order. Unknowns stay listed.6AnswerRanked where the evidence separates models, and flagged where it can't.1 · Where evidence comes fromA value with no source never reaches a decision. Neither does one that only its collector has checked.
Model cards hold the measurements, one record per result. Benchmark pages say what each benchmark measures, and which domains it counts toward.
Every measurement carries one of three labels. A lab's claim about its own model sits beside independent results, labelled, never mixed in unmarked.
benchmark_authorThe benchmark's own authors.independent_evaluatorAn evaluator outside the lab.provider_self_reportThe lab or provider, about its own model.A record enters the snapshot only when a second check agrees with the first, and that check was made by a deterministic tool or by a model from a different family.
A value whose latest check isn't "verified", anything without a registered source URL, anything from an excluded source, or the old flat score block on each card.
Two publishers whose terms don't permit our use were removed on 24 September 2026, with every value and benchmark they own. A test fails the build if either comes back, and the snapshot build scans its own output with the same rules.
Check it
models/the model cardsbenchmarks/benchmark pages and domain tagsschema/card.pysource_kind: the three measurer labelsdecision/model.pysources required; the two-key ruletests/test_removed_sources.pykeeps excluded sources outevidence_basis meansThe rank export and the /v1/rank API label each row's inputs with evidence_basis. The label describes the inputs. It is not a verdict on the model.
noneNo usable measurement contributed.unverified-legacyEvery input is an older card value.mixedReviewed and older values both contribute.partial-verifiedEvery input is reviewed, but some the profile asks for are missing.verifiedEvery weighted benchmark is present and reviewed.Check it
pipeline/ranking.py_basis(): the evidence_basis labels2 · The capability estimateThere is no fixed benchmark list and no hand-set weight. Every admitted benchmark with at least two model observations counts, and nobody picks favourites. One leaderboard is one reading, and readings disagree. The estimate uses all of them, and says how sure it is.
The fit reads which domains a benchmark counts toward from that benchmark's own page, tagged direct or proxy.
A direct measurement loads at 1, a proxy at 0.35, so a proxy carries 0.1225 of a direct one's precision.
A measurement's precision halves every 365 days, down to a floor of a quarter.
A central 80% range: the estimate ± 1.2816 standard deviations.
A model with nothing measured in a domain gets no estimate there.
The weights don't depend on the scores.
Check it
decision/capability.pyfit_capabilities(): the fit itselfdocs/design/capability-model.mdwhy this model, and its rulesdocs/validation/capability-model.mdthe held-out test report3 · TiesEvery estimate is a range. When two models' ranges overlap, the evidence can't say which is better, so the page doesn't pretend to.
range = estimate ± 1.2816 × sd
not_separable = max(A.low, B.low) ≤ min(A.high, B.high)
for any other model BA row is flagged not_separable when its range overlaps the range of any other model in the domain being ranked, not only the leader's. Two models can be told apart from the leader and still not from each other. The board names the models the evidence cannot separate. The API keeps a stable transport order so a list can travel. Read the flag before the rank.
Ability cannot split a tied group. Price, context, licence and the other facets you set can, and the board shows them for every row.
Check it
decision/engine.pywhere not_separable is setdecision/capability.pydeterministic_probabilities()docs/decision-contract.mdp_best and top3_stability4 · Must and PreferEvery facet on the board has three settings, and each does a different job. Conditions filter. They never add points.
The default. The row stays on the board, greyed, so you can always see what you didn't choose.
A hard gate, with a threshold where the facet has one: context of at least 200K, a cost cap, a signed BAA. A model that fails is excluded, and stays on screen as excluded.
A weight. It orders the models that passed every Must. It never removes one.
For agents, in a spec. Musts are the conditions under where. Prefers are the weights under optimize. A condition marked soft(penalty) is a preference too.
Check it
docs/design/briefs/2026-09-27-facet-board.mdthe three settings, and whydecision/filter.pyno condition adds a scoredocs/decision-contract.mdsoft conditions5 · Unknown means unknownEvery condition has three answers: pass, fail and unknown. Unknown is its own answer, with its own rules. A null beats a guess.
| The value is | Which means | What the answer does |
|---|---|---|
| Known | A value with a source, and a check that passed. | Used. |
| Unknown, on ability or a spec such as context | Nobody has published it, or it hasn't been collected yet. | May qualify: listed beside the answer, not ranked, never scored as zero. |
| Unknown, on your data or the licence | We can't confirm the provider's terms. | Not treated as met. Shown as unverified, may qualify, so you can check it yourself. |
| Inapplicable | The model's class can't have it. A model that writes no text has no output-token limit. | Derived from the class, never typed on a card. |
Check it
decision/filter.pypass, fail and unknowndocs/decision-contract.mdwhat an unknown doesschema/applicability.pyinapplicable, derived from the class6 · ReproducibilityA decision depends on two things you can name: the question and the evidence. Pin both, and anyone holding the same snapshot file gets the answer you got.
Every admitted record, written as canonical JSON: sorted keys, no whitespace.
Key order, whitespace and compact or YAML spelling don't change the hash. Changing any field does.
dec_ plus a hash of the spec hash and the snapshot id.
This page reads the signature block from the decision snapshot that the site publishes.
$ modelspec snapshot fetch $ modelspec decide spec.yaml --json "snapshot": "[SNAPSHOT ID]", "spec_hash": "sha256:[SPEC HASH]", "signature_verified": true
This snapshot is Ed25519-signed with published key ID ed25519-dd3ae9b0899184e5. The signature was verified when this page was built, with the same check the CLI runs.
Run modelspec snapshot fetch to download the snapshot and verify its signature yourself.
The content hash is checked separately.
Check it
docs/decision-snapshot.mdthe file format, hash and iddocs/decision-contract.mdthe canonical spec hashdocs/snapshot-signing.mdhow snapshots are signeddecision/snapshot_keys.jsonthe public keys the CLI pins7 · NeutralityWe charge the people and agents who ask for an answer. Never the models, labs or hosts the answer is about.
“No referral fees, no paid placement, no provider-paid visibility, permanently.”
accepts_referral_feesfalseaccepts_paid_placementfalseaccepts_provider_paid_visibilityfalseproxies_inference_tokensfalsestores_customer_promptsfalseconceals_purchases_from_catalogued_vendorsfalselets_supplier_models_write_supplier_cardsfalseCheck it
docs/legal/neutrality.mdthe commitment, in fullapi/ranking/engine.pyneutrality_commitment()8 · What we don't doYou set each criterion yourself. No parser and no AI reads your words and guesses what you meant.
No provider, host or gateway can buy inclusion, position or a mention.
We recommend and hand off. Your inference goes to the provider directly.
A request carries a profile, not prompt text, so there is nothing to keep.
A missing value stays missing. It is never filled in, and never counted as zero.
Nobody chooses which benchmarks matter. Every admitted benchmark counts, weighted by the same rules.
The decision snapshot and vocabulary read by decision answers are public, versioned and free to fetch. So is the code that reads them. The /v1/policy-check determinations are private.
modelspec.dev/api/decision/snapshot.json.gzThe snapshot that decision answers read.modelspec.dev/api/decision/vocabulary.jsonEvery facet, benchmark and domain decision answers know.modelspec.dev/api/rank/profiles.jsonThe ranking floors and the neutrality commitment, as data.modelspec.dev/.well-known/modelspec-snapshot-keys.jsonThe public keys used to verify signed snapshots.