DeepSeek V4 Flash Slashes AI Reasoning Costs by 99% in 19 Months

ARC Prize's latest chart shows collapsing per-task costs on ARC-AGI benchmarks, with reasoning models now solving problems for pennies instead of tens of dollars.

·
·
DeepSeek V4 Flash Slashes AI Reasoning Costs by 99% in 19 Months
Read5 min
TypeNews
  • ARC-AGI-1 cost to reach 75% fell 99.95% in 19 months, from $26 to $0.01 per task.
  • ARC-AGI-2 saw a 99.44% cost drop in six months, from $13.62 to $0.08 per task.
  • DeepSeek V4 Flash scores 89.0% on ARC-AGI-1 and 61.4% on ARC-AGI-2 at max effort.
  • Gemini 3 Deep Think still leads ARC-AGI-2 at 84.6% accuracy but costs $13.62 per task.
  • Harness scaffolding alone can triple benchmark scores without changing the underlying model.
  • Agentic workflows using 50 reasoning calls now cost around $1 instead of $250.

ARC-AGI reasoning costs fall from dollars to cents

The latest ARC Prize leaderboard snapshot shows the cost of solving abstract-reasoning tasks falling by more than 99% over periods measured in months. Lower inference prices, sparse model architectures, adaptive reasoning and better evaluation software all contribute. For developers, the results suggest that repeated reasoning and verification can support wider agent workflows, provided the economics hold on application-specific tests.

ARC tests rule-finding from sparse examples

François Chollet created ARC-AGI to test whether a system can infer unfamiliar rules from a few examples. Each task presents input and output grids, then asks the system to apply the inferred transformation to a new grid. The novel visual patterns reduce the value of memorizing training data and make each task a compact test of abstraction.

ARC-AGI-2 increases the difficulty through tasks that require several interacting concepts and more precise transformations. ARC Prize’s results analysis reports that pure LLM baselines scored 0% under its evaluation setup, while human participants could solve every task.

The leaderboard publishes accuracy alongside estimated cost per task. That cost represents API spend under a submitted harness, including its prompts, retries and generated tokens. Hardware expense, latency and engineering overhead sit outside the figure.

A penny reaches the ARC-AGI-1 line

Cost reductions reported in the cited leaderboard snapshot
Benchmark Comparison Earlier cost Newer cost Interval Reduction
ARC-AGI-1 75% reference line $26.00 per task $0.01 per task 19 months 99.96%
ARC-AGI-2 Charted cost comparison $13.62 per task $0.08 per task 6 months 99.41%

DeepSeek V4 Flash 0731 reaches the ARC-AGI-1 chart’s 75% reference line for about one cent per task, down from roughly $26 for o3-preview 19 months earlier. The ARC-AGI-2 row uses a separate charted comparison because DeepSeek’s maximum-effort score on that benchmark remains below 75%.

DeepSeek V4 Flash at maximum effort on semi-private sets
Benchmark Accuracy Cost per task
ARC-AGI-1 Semi-Private 89.0% $0.02
ARC-AGI-2 Semi-Private 61.4% $0.04

The one-cent ARC-AGI-1 result uses the cheaper effort setting needed to cross 75%. At maximum effort, the model spends about two cents per task and reaches 89.0%. Effort settings therefore need to accompany every cost comparison.

Top scores still carry a premium

Selected high-scoring leaderboard entries
Benchmark Model Accuracy Cost per task
ARC-AGI-1 Gemini 3 Deep Think Above 96% $7.17
ARC-AGI-1 Opus 4.6 93.0% $1.88
ARC-AGI-1 GPT-5.2 Pro 90.5% $11.64
ARC-AGI-2 Gemini 3 Deep Think 84.6% $13.62

Gemini 3 Deep Think leads the cited ARC-AGI-2 results with 84.6% accuracy at $13.62 per task. On ARC-AGI-1, GPT-5.2 Pro’s $11.64 cost is about 387 times lower than o3’s reported $4,500 per task one year earlier, although the runs use different models, scores and evaluation configurations.

The spread between DeepSeek’s low-cost results and the highest scores exposes a practical trade-off. Small accuracy gains near the top can require substantially more inference, while many applications may obtain sufficient performance from a cheaper effort level.

Four levers push costs down

  1. Sparse activation reduces computation. DeepSeek’s 552-billion-parameter mixture-of-experts architecture reportedly activates about 8 billion parameters while processing the prompt and 16 billion while generating the answer. Most parameters remain inactive for each token, lowering the work required to serve the model.
  2. Adaptive reasoning limits token use. In the cited DeepSeek comparison, the maximum-effort setting consumed fewer reasoning tokens overall than the high-effort setting while solving more tasks. The result suggests that improved stopping behavior can raise accuracy without increasing aggregate token use.
  3. Harness design changes outcomes. A harness formats examples, calls the model, parses its output and manages retries. An OpenAI experiment increased an ARC-AGI score from about 13% to 40% by replacing the public harness with an internally optimized version while keeping the model fixed.
  4. Provider pricing lowers measured spend. DeepSeek made a promotional 75% discount permanent, reducing V4 Pro pricing from a range of $0.0145 to $3.48 per million tokens to $0.003625 to $0.87. Because leaderboard costs use provider prices, commercial discounts can reduce dollars per task even when token consumption stays constant.

Production budgets need their own math

Because each ARC task has a specific prompt, output format, retry policy and harness, its cost cannot be copied directly into a coding-agent budget. A mechanical comparison illustrates the scale: 50 ARC-like tasks at $0.02 each would cost $1, while 50 tasks at $11.64 each would cost $582. Real agent calls may use longer contexts, tools, cached tokens and multiple validation passes.

Teams evaluating these models should measure the full cost of a successful application outcome through a controlled workload:

  • Reproduce the workflow. Use production prompts, tools, parsers and retry rules instead of relying on raw benchmark scores.
  • Track accepted results. Divide total spend by outputs that pass application-specific checks, including failed attempts and verification calls.
  • Tune effort dynamically. Route simple requests to cheaper settings and escalate uncertain cases to deeper reasoning.
  • Measure operational performance. Record median and tail latency, rate-limit failures, context usage and output variance.
  • Check deployment constraints. Confirm API availability, regional support, privacy terms and service limits, which the ARC leaderboard does not cover.

The snapshot has clear limits

Semi-private evaluations are distinct from the private set used for ARC Prize eligibility, and harness choices can materially affect both score and cost. ARC tasks also focus on compact grid transformations, leaving coding, tool use, long-context retrieval and domain knowledge untested.

Repeated reasoning is becoming cheap enough for broader search, verification and agent fan-out. Application-specific evaluations must confirm that the lower cost survives real prompts, production constraints and the required success rate.

Trending
  • No trending articles

Comments

avatar

Next Reads