DeepSeek V4 Flash 0731 scores 89% on ARC-AGI-1
Original: DeepSeek V4 Flash 0731
Why This Matters
High ARC-AGI-2 scores at sub-cent costs signal rapid progress in efficient reasoning AI models.
DeepSeek released V4 Flash 0731 on July 31, 2026, achieving 89.0% on ARC-AGI-1 Semi-Private and 61.4% on ARC-AGI-2 Semi-Private at max effort, with costs as low as $0.02 and $0.04 per task respectively, across three reasoning variants.
DeepSeek released V4 Flash 0731 on July 31, 2026, submitting results to the ARC Prize leaderboard with three reasoning variants: Max, High, and Low. At maximum effort, the model scored 89.0% on ARC-AGI-1 Semi-Private at $0.02 per task, and 61.4% on ARC-AGI-2 Semi-Private at $0.04 per task. The High variant scored 87.0% and 56.0%, while the Low variant scored 84.0% and 46.0% on ARC-AGI-1 and ARC-AGI-2 respectively. ARC-AGI-2 is widely regarded as a significantly harder benchmark than ARC-AGI-1. The results are verified by ARC Prize and include pass/fail data across 120 ARC-AGI-2 public eval tasks and 400 ARC-AGI-1 public eval tasks. No ARC-AGI-3 scores were reported. The model's combination of high benchmark performance and low per-task cost positions it as a competitive option in the reasoning model space.