Model record
Claude Opus 4.5
reasoningAnthropic · released Nov 24, 2025
Overall standing
- Overall
- 48.3rank 34 of 63
- 48.3 is at approximately the 47th percentile of the ranked Overall field. The field spans 19.0 to 62.2, with a median of 49.4.
- 6 of 8 measured · 2 estimated
Category standingsReasoning 61.7Coding 57.2Math 4.6Agentic 20.3
- Reasoning
- 61.7rank 40 of 63
- 61.7 is at approximately the 37th percentile of the ranked Reasoning field. The field spans 31.6 to 73.3, with a median of 65.0.
- 4 of 4 measured
- Coding
- 57.2rank 27 of 63
- 57.2 is at approximately the 58th percentile of the ranked Coding field. The field spans 10.4 to 72.4, with a median of 54.5.
- 1 of 2 measured · 1 estimated
- Math
- 4.6rank 34 of 63
- 4.6 is at approximately the 46th percentile of the ranked Math field. The field spans 0.0 to 32.3, with a median of 5.7.
- 1 of 1 measured
- Agentic
- 20.3rank 29 of 63
- 20.3 is at approximately the 55th percentile of the ranked Agentic field. The field spans 3.3 to 33.4, with a median of 19.0.
- 0 of 1 measured · 1 estimated
Model factsProprietary · $5.0 / $25.0 per 1M
Model facts come from the provider page. Listed API prices carry a separate first-party source and check date.
Scores and sources
Measured scores retain their source and retrieval date. Missing scores are not treated as zero.
| Benchmark | Category | Score | Evaluation settings | Source |
|---|---|---|---|---|
| Terminal-Bench v2.1 | Coding | — | Not measured in this dataset | No source |
| τ³-Banking | Agentic | — | Not measured in this dataset | No source |
| AA-LCR | Reasoning | 74.0 | Artificial Analysis independent run; 100 roughly 100k-token multi-document questions; 3 repeats; equality-checker pass@1; no tools | Open source (opens in a new tab) |
| Humanity's Last Exam | Reasoning | 28.4 | Artificial Analysis independent run; 2,158 text-only questions; pass@1; GPT-4o (August 2024) equality checker; no tools | Open source (opens in a new tab) |
| GPQA Diamond | Reasoning | 86.6 | Artificial Analysis independent run; 198 questions; 5 repeats; regex-graded pass@1; no tools | Open source (opens in a new tab) |
| SciCode | Coding | 49.5 | Artificial Analysis independent run; 288 test subproblems; scientist background included; 3 repeats; subproblem pass@1 | Open source (opens in a new tab) |
| IFBench | Reasoning | 58.0 | Artificial Analysis independent run; 294 single-turn prompts; 5 repeats; official loose evaluator; prompt-level pass@1 | Open source (opens in a new tab) |
| CritPt | Math | 4.6 | Artificial Analysis independent run; 70 test challenges; 5 repeats; two-step answer parsing; official CritPt grader; pass@1; no tools | Open source (opens in a new tab) |