Methodology

LM Board runs no evaluations. It records benchmark scores published by their named source — today, all 463measured scores are published by Artificial Analysis — then computes one equal-weight Index per category. Every measured score keeps a source link, retrieval date, and any available evaluation settings. Missing benchmark results are estimated only inside the Index and are disclosed as estimates. An Overall rank requires 60% measured coverage; clearing that broad evidence gate also permits complete estimated category Indexes.

Where the scores come from

All 463 scores on the board carry a source link and the date they were retrieved, and every one records the settings it was run under. The link and the date are required by the data schema, so a result that arrives without both is rejected at build time rather than published unsourced. Open any model record from the leaderboard to read them.

Third-party measurements are preferred over a lab's own reporting. Today none of the 463scores on the board are self-reported — every one is an Artificial Analysis (opens in a new tab) measurement. A score that did come from the model's maker would stay on the board and carry a Vendor mark next to the number.

Price, context window and release date are not benchmark results and are not sourced the same way: they come from the provider's own public listing, linked as Official page on every model record, and carry no separate retrieval date. 13 of the 63 models on the board publish no price at all, and show a dash rather than a zero.

How the Index is calculated

A model's Index is the plain average of its scores across every benchmark on the tab. Every benchmark counts equally — no weighting, no Elo, no adjustments.

Index = sum of a model's scores ÷ number of benchmarks on the tab

Where a model has no result, the average uses an estimate rather than skipping the benchmark: the model's standing on the benchmarks it wasmeasured on, read off the missing benchmark's own spread of published results. A model that ranks mid-field elsewhere is credited a mid-field result, so skipping a benchmark neither helps nor hurts. A missing score is never counted as zero, and an estimate is never published as a score — the table still shows “—” in that column. A category Index may be entirely estimated when a broadly measured model has no result in that category; it is labeled as estimated rather than presented as a measurement.

Example of how the Index and ranks behave with missing scores
RankModelIndexBench 1Bench 2Bench 3Bench 4
1Model B87.892.093.0 estimated88.078.0
2Model A81.080.090.070.084.0
UnrankedModel CInsufficient data95.096.0
An illustrative Overall suite with four benchmarks, where ranking requires three. Model B ranks first on the strength of the three benchmarks it was measured on; its Bench 2 gap is estimated at the level it performs elsewhere, so the omission is neither a penalty nor a free pass. Model C scores well but covers only two of four benchmarks, so it keeps its scores and gets no rank — and no estimates.

Each tab — Overall, Reasoning, Coding, Math, and Agentic — applies the same average to its own set of benchmarks, and the tabs are not the same size. Where a category has a single benchmark — Math and Agentic today — the Index is that benchmark's measured score when available, or a disclosed estimate for a model that cleared the Overall evidence gate.

Who gets ranked: the 60% rule

Broad evidence comes first. An Overall rank requires measured scores on at least 60% of the full suite — currently 5 of 8. Once a model clears that gate, every category can receive a complete Index: measured category scores are used directly and every remaining gap is estimated from the model's standing across its measured benchmarks. A model that has not cleared Overall can still rank in a category by measuring at least 60% of that category.

If a model clears neither route, it still appears with every measured score but shows “Insufficient data” and no rank. Estimates never become score records, and an incomplete estimated Index is rejected. This keeps sparse evidence from becoming a free pass while allowing broadly tested models to be compared across every category.

Models with the same Index share the same rank, and the next distinct Index skips the ranks the tie used up — two models tied at 2nd are followed by a 4th, not a 3rd. Nothing outside the scores breaks a tie.

Search and filters never change the numbers: they only hide rows, so a model keeps the same rank however the table is narrowed.

The benchmarks

The board currently tracks 8benchmarks across four categories. All of them report scores on a 0–100 scale, which is what makes a direct average possible; a benchmark on a different scale would still be displayed, but would stay out of the Index.

They do not share a difficulty, though, so a bar drawn as a fraction of 100 would compare benchmarks rather than models. The bar under each score instead shows where that score falls within the range the benchmark has actually produced across every model on the board.

Reasoning

Coding

  • Terminal-Bench v2.1Verified terminal-agent tasks spanning software engineering, system administration, data processing, training, and security.Source (opens in a new tab)
  • SciCodeScientist-curated Python coding problems drawn from realistic research workflows across 16 subfields.Source (opens in a new tab)

Math

  • CritPtQuantitative research-level physics challenges with numerical, symbolic, and executable-function answers.Source (opens in a new tab)

Agentic

  • τ³-BankingKnowledge-intensive banking support workflows requiring policy retrieval and multi-step tool-mediated state changes.Source (opens in a new tab)

Honest limits

Published scores can use different tools, prompting setups, and reasoning budgets. The board does not publish confidence intervals, so small gaps should not be treated as proof of a meaningful capability difference. When a model row shows a reasoning-effort label, that setting applies to every score in the row.

Results and provider pricing change as labs publish updates; the linked sources remain authoritative.

Corrections welcome; the issue tracker will be linked at publish time.