Open methodology · v1

Opinion, with the workings shown.

This is not an objective benchmark. It is a legible summary of how the community feels about the models it has actually used.

01 / BALLOTS

Rank what you know

Members place models from S through F and may leave unfamiliar ones unranked. At least five placements are required to publish. There is one current ballot per person and board; editing replaces the previous ballot’s influence while preserving revision history.

02 / SCORE

Abstention is not a zero

S through F map to 6 through 1. For each model we first take the observed mean only from ballots that ranked it. Then we pull uncertain results toward the category baseline using ten equivalent ballots.

score = (votes × observed mean + 10 × board baseline) / (votes + 10)

This keeps a brand-new release with two enthusiastic voters from appearing more certain than a model judged by hundreds.

03 / TIERS

Stable bands, honest detail

Fixed score bands determine the visible tier. Within each tier, models are ordered by score and then unique voter count. Every model page shows its voter count and complete S–F distribution, so disagreement stays visible.

04 / FRESHNESS

Current without the theatre

Public results may be cached for up to five minutes. Production will publish precomputed snapshots rather than recalculating raw ballots on every visit. Algorithm changes get a new version and changelog.

LOCAL DEMO

Everything stays in this browser

This implementation uses seeded model facts and browser storage for ballots, comments, proposals, votes, and shared snapshots. It makes no catalog, identity, database, email, or analytics requests.