Model record
DeepSeek V4 Flash 0731
Open weightsDeepSeek · released Jul 31, 2026
Overall standing
- Overall
- Insufficient data
- 0 of 8 measured
Category standingsReasoning —Coding —Math —Agentic —
- Reasoning
- Insufficient data
- 0 of 4 measured
- Coding
- Insufficient data
- 0 of 2 measured
- Math
- Insufficient data
- 0 of 1 measured
- Agentic
- Insufficient data
- 0 of 1 measured
Model factsOpen weights · $0.14 / $0.28 per 1M
Copied from the linked provider page; no separate retrieval date is stored.
Scores and sources
Measured scores retain their source and retrieval date. Missing scores are not treated as zero.
| Benchmark | Category | Score | Evaluation settings | Source |
|---|---|---|---|---|
| Terminal-Bench v2.1 | Coding | — | Not measured in this dataset | No source |
| τ³-Banking | Agentic | — | Not measured in this dataset | No source |
| AA-LCR | Reasoning | — | Not measured in this dataset | No source |
| Humanity's Last Exam | Reasoning | — | Not measured in this dataset | No source |
| GPQA Diamond | Reasoning | — | Not measured in this dataset | No source |
| SciCode | Coding | — | Not measured in this dataset | No source |
| IFBench | Reasoning | — | Not measured in this dataset | No source |
| CritPt | Math | — | Not measured in this dataset | No source |