Model record
Mistral Medium 3.5
Open weightsMistral AI · released Apr 28, 2026
Overall standing
- Overall
- 40.2rank 49 of 62
- 40.2 is at approximately the 22th percentile of the ranked Overall field. The field spans 19.0 to 62.2, with a median of 48.9.
- 8 of 8 measured
Category standingsReasoning 54.4Coding 45.1Math 0.0Agentic 14.4
- Reasoning
- 54.4rank 51 of 62
- 54.4 is at approximately the 19th percentile of the ranked Reasoning field. The field spans 31.6 to 73.3, with a median of 64.8.
- 4 of 4 measured
- Coding
- 45.1rank 36 of 46
- 45.1 is at approximately the 23th percentile of the ranked Coding field. The field spans 10.4 to 72.4, with a median of 56.1.
- 2 of 2 measured
- Math
- 0.0rank 53 of 62
- 0.0 is at approximately the 8th percentile of the ranked Math field. The field spans 0.0 to 32.3, with a median of 5.3.
- 1 of 1 measured
- Agentic
- 14.4rank 31 of 46
- 14.4 is at approximately the 34th percentile of the ranked Agentic field. The field spans 3.3 to 33.4, with a median of 19.6.
- 1 of 1 measured
Model factsOpen weights · $1.5 / $7.5 per 1M
Copied from the linked provider page; no separate retrieval date is stored.
Scores and sources
Measured scores retain their source and retrieval date. Missing scores are not treated as zero.
| Benchmark | Category | Score | Evaluation settings | Source |
|---|---|---|---|---|
| Terminal-Bench v2.1 | Coding | 50.6 | Artificial Analysis independent run; 89 tasks; Terminus 2 on E2B; 3 repeats; pass@1; 250-episode cap; 2-hour timeout | Open source (opens in a new tab) |
| τ³-Banking | Agentic | 14.4 | Artificial Analysis independent run; 97 tasks; 5 repeats; backend-state pass@1; BM25 plus grep retrieval; 200-step cap | Open source (opens in a new tab) |
| AA-LCR | Reasoning | 61.0 | Artificial Analysis independent run; 100 roughly 100k-token multi-document questions; 3 repeats; equality-checker pass@1; no tools | Open source (opens in a new tab) |
| Humanity's Last Exam | Reasoning | 12.8 | Artificial Analysis independent run; 2,158 text-only questions; pass@1; GPT-4o (August 2024) equality checker; no tools | Open source (opens in a new tab) |
| GPQA Diamond | Reasoning | 74.9 | Artificial Analysis independent run; 198 questions; 5 repeats; regex-graded pass@1; no tools | Open source (opens in a new tab) |
| SciCode | Coding | 39.6 | Artificial Analysis independent run; 288 test subproblems; scientist background included; 3 repeats; subproblem pass@1 | Open source (opens in a new tab) |
| IFBench | Reasoning | 68.8 | Artificial Analysis independent run; 294 single-turn prompts; 5 repeats; official loose evaluator; prompt-level pass@1 | Open source (opens in a new tab) |
| CritPt | Math | 0.0 | Artificial Analysis independent run; 70 test challenges; 5 repeats; two-step answer parsing; official CritPt grader; pass@1; no tools | Open source (opens in a new tab) |