DeepSeek V4 Flash 0731 scores 89.0% on ARC-AGI-1, 61.4% on ARC-AGI-2

DeepSeek released DeepSeek V4 Flash 0731 on July 31, 2026, in three reasoning-effort variants: Max, High and Low. ARC Prize, which independently verifies model scores on its ARC-AGI benchmark series rather than taking vendor-reported numbers, published results for all three on its leaderboard. At Max effort the model scores 89.0% on ARC-AGI-1 Semi-Private at a cost of $0.02 per task, and 61.4% on ARC-AGI-2 Semi-Private at $0.04 per task. At High effort it scores 87.0% on ARC-AGI-1 and 56.0% on ARC-AGI-2. At Low effort it scores 84.0% on ARC-AGI-1 and 46.0% on ARC-AGI-2. ARC Prize gives a per-task cost only for the Max variant, not for High or Low. The results table shows a dash for ARC-AGI-3 across all three variants, so no score for that benchmark is reported. The page also shows a pass/fail grid across the ARC-AGI-2 Public Eval set, which contains 120 tasks, but includes no model architecture, parameter count or training details, and no comparison to other models' scores or rankings.
Key facts
- DeepSeek V4 Flash 0731, released July 31, 2026, comes in three reasoning-effort variants: Max, High and Low.
- At Max effort it scores 89.0% on ARC-AGI-1 Semi-Private at $0.02 per task and 61.4% on ARC-AGI-2 Semi-Private at $0.04 per task.
- At High effort it scores 87.0% on ARC-AGI-1 and 56.0% on ARC-AGI-2; at Low effort, 84.0% and 46.0%.
- ARC Prize publishes a per-task cost only for the Max variant, and no ARC-AGI-3 score is given for any of the three.
- The ARC-AGI-2 Public Eval set used for the pass/fail grid contains 120 tasks.
Why it matters
ARC-AGI is built to test novel reasoning rather than pattern memorization, and ARC Prize verifies scores itself rather than publishing vendor claims. A verified result for a new DeepSeek model, at a few cents per task, adds a public, third-party data point to the ARC-AGI leaderboard the same day the model appeared.
Who it affects
Anyone choosing a model for reasoning-heavy tasks, evaluating DeepSeek's lineup specifically, or tracking the ARC-AGI series as a general proxy for reasoning progress across labs.
How to use it
The three variants trade accuracy for reasoning effort. Max gets the highest scores, 89.0% on ARC-AGI-1 and 61.4% on ARC-AGI-2, at $0.02 and $0.04 per task respectively. High effort scores 87.0% and 56.0%; Low effort scores 84.0% and 46.0%. ARC Prize does not publish a per-task cost for the High or Low variants, only for Max, so a workload's variant choice can weigh the accuracy gap, roughly 2 to 5 percentage points between Max and High and 3 to 10 points between High and Low, but not the price difference.
How solid is it
The scores come from ARC Prize's verified leaderboard, meaning the organization ran the benchmark itself rather than relying on DeepSeek's own reported figures. The results page gives no model architecture, parameter count or training details, and no comparison to other models' scores or rankings, so the numbers stand without ranking context from the source itself.
Risks and caveats
The table lists a dash for ARC-AGI-3 across all three variants, so there is no data yet on the newest benchmark in the series. Per-task cost is stated only for the Max variant; the cost of running High or Low effort is unstated. The page also does not explain the difference between the Semi-Private set used for the headline scores and the Public Eval set of 120 tasks shown in the pass/fail grid, so the two score sets cannot be cross-checked against each other from the page alone.