Methodology

LM Board runs no evaluations. It records benchmark scores published by their named source — today, all 456 measured scores are published by Artificial Analysis — then computes one equal-weight Index per category. Every measured score keeps a source link, retrieval date, and any available evaluation settings. Missing benchmark results are estimated only inside the Index and are disclosed as estimates; models below 60% measured coverage are not ranked.

Where the scores come from

All 456 scores on the board carry a source link and the date they were retrieved, and every one records the settings it was run under. The link and the date are required by the data schema, so a result that arrives without both is rejected at build time rather than published unsourced. Open any model record from the leaderboard to read them.

Third-party measurements are preferred over a lab's own reporting. Today none of the 456 scores on the board are self-reported — every one is an Artificial Analysis (opens in a new tab) measurement. A score that did come from the model's maker would stay on the board and carry a Vendor mark next to the number.

Price, context window and release date are not benchmark results and are not sourced the same way: they come from the provider's own public listing, linked as Official page on every model record, and carry no separate retrieval date. 12 of the 62 models on the board publish no price at all, and show a dash rather than a zero.

How the Index is calculated

A model's Index is the plain average of its scores across every benchmark on the tab. Every benchmark counts equally — no weighting, no Elo, no adjustments.

Index = sum of a model's scores ÷ number of benchmarks on the tab

Where a model has no result, the average uses an estimate rather than skipping the benchmark: the model's standing on the benchmarks it was measured on, read off the missing benchmark's own spread of published results. A model that ranks mid-field elsewhere is credited a mid-field result, so skipping a benchmark neither helps nor hurts. A missing score is never counted as zero, and an estimate is never published as a score — the table still shows “—” in that column, and the model row reports how many of its benchmarks were estimated.

Example of how the Index and ranks behave with missing scores
RankModelIndexBench 1Bench 2Bench 3Bench 4
1Model B87.892.093.0 estimated88.078.0
2Model A81.080.090.070.084.0
UnrankedModel CInsufficient data95.096.0
An illustrative tab with four benchmarks, where ranking requires three. Model B ranks first on the strength of the three benchmarks it was measured on; its Bench 2 gap is estimated at the level it performs elsewhere, so the omission is neither a penalty nor a free pass. Model C scores well but covers only two of four benchmarks, so it keeps its scores and gets no rank — and no estimates.

Each tab — Overall, Reasoning, Coding, Math, and Agentic — applies the same average to its own set of benchmarks, and the tabs are not the same size. Where a category has a single benchmark — Math and Agentic today — the Index is that benchmark's score, and the word "average" is doing no work.

Who gets ranked: the 60% rule

An average over two benchmarks says less than an average over eight, so a model is ranked only once it has scores on at least 60% of a tab's benchmarks. On the Overall tab that is currently 5 of 8. Below that bar a model still appears with every score it has, but shows “Insufficient data” in place of an Index and “—” in place of a rank. Only measured results count toward the bar — estimates fill the gaps of a model that already cleared it, never carry a model over it. Without this rule, a model evaluated only on its strongest few benchmarks could top the table.

Models with the same Index share the same rank, and the next distinct Index skips the ranks the tie used up — two models tied at 2nd are followed by a 4th, not a 3rd. Nothing outside the scores breaks a tie.

Search and filters never change the numbers: they only hide rows, so a model keeps the same rank however the table is narrowed.

The benchmarks

The board currently tracks 8 benchmarks across four categories. All of them report scores on a 0–100 scale, which is what makes a direct average possible; a benchmark on a different scale would still be displayed, but would stay out of the Index.

They do not share a difficulty, though, so a bar drawn as a fraction of 100 would compare benchmarks rather than models. The bar under each score instead shows where that score falls within the range the benchmark has actually produced across every model on the board.

Reasoning

Coding

  • Terminal-Bench v2.1Verified terminal-agent tasks spanning software engineering, system administration, data processing, training, and security.Source (opens in a new tab)
  • SciCodeScientist-curated Python coding problems drawn from realistic research workflows across 16 subfields.Source (opens in a new tab)

Math

  • CritPtQuantitative research-level physics challenges with numerical, symbolic, and executable-function answers.Source (opens in a new tab)

Agentic

  • τ³-BankingKnowledge-intensive banking support workflows requiring policy retrieval and multi-step tool-mediated state changes.Source (opens in a new tab)

Honest limits

Published scores can use different tools, prompting setups, and reasoning budgets. The board does not publish confidence intervals, so small gaps should not be treated as proof of a meaningful capability difference. When a model row shows a reasoning-effort label, that setting applies to every score in the row.

Results and provider pricing change as labs publish updates; the linked sources remain authoritative.

Corrections welcome; the issue tracker will be linked at publish time.