Google DeepMind's Gemini 3.7 Flash Doubles Coding Scores at Half the Price

Gemini 3.7 Flash lands just three weeks after 3.6, with major coding and agent gains at half the price

·
·
Google DeepMind's Gemini 3.7 Flash Doubles Coding Scores at Half the Price
Read5 min
TopicLlms · Api
  • Gemini 3.7 Flash is live — available now in Google AI Studio, Antigravity, Android Studio, and the Gemini app.
  • Introductory pricing is half of 3.6 Flash: $0.75/1M input and $3.75/1M output tokens, through end of 2026.
  • Major coding gains: DeepSWE v1.1 jumps from 49.0% (3.6 Flash) to 65.3%; FrontierCode goes from 34.4% to 43.6%.
  • Knowledge work nearly doubles: AutomationBench score goes from 17.0% to 30.4% over 3.6 Flash.
  • Gemini Spark upgraded: The 24/7 personal agent for Pro/Ultra subscribers now runs on 3.7 Flash with better Workspace tool use.
  • Released just 3 weeks after 3.6 Flash, signaling Google's aggressive iterative cadence on the Flash model line.

Google DeepMind just shipped Gemini 3.7 Flash, the newest entry in its workhorse Flash series. The headline is not just the performance jump , it is the pace. This release comes just three weeks after Gemini 3.6 Flash, and is a direct result of developer feedback and algorithmic innovations that Google plans to bring to future models. At this cadence, the Flash line is evolving faster than most teams can finish integrating the previous version.

3.7 Flash delivers substantial improvements across software engineering, knowledge work, and web development workflows , with an introductory price of half the original 3.6 Flash cost per million tokens. That price-to-performance shift is the real story here.

The numbers that matter

The benchmark gains over 3.6 Flash are significant, especially on the coding side. Google measured 3.7 Flash on two key software engineering evals:

  • FrontierCode 1.1 Main: 43.6% vs 34.4% for 3.6 Flash , a measure of production-ready code quality.
  • DeepSWE v1.1: 65.3% vs 49.0% , this benchmark tests whether a model can autonomously resolve real software issues end-to-end.

Web development also sees a meaningful jump. 3.7 Flash generates more functional layouts and feature-complete apps in fewer prompts, and outperforms 3.6 Flash on Arena.ai's WebDev Arena with an Elo score of 1588 vs 1538. For context, Elo scores here work like chess ratings , a 50-point gap at this level is a meaningful, consistent advantage.

Knowledge-dense professional domains get a particularly large upgrade. 3.7 Flash significantly outperforms 3.6 Flash on the GDP.pdf benchmark (34.0% vs 22.0%), an eval for testing a model's ability to process complex documents. It also surpasses 3.6 Flash in AutomationBench, demonstrating it can more effectively complete real-world business workflows (30.4% vs 17.0%). That AutomationBench gap , nearly double the score , is the kind of jump that changes whether you trust a model to run unsupervised in a production pipeline.

Smarter agent behavior, not just better scores

Beyond raw benchmarks, Google is emphasizing a qualitative shift in how 3.7 Flash behaves as an agent. It better adapts to roadblocks, clarifies intent when needed, and follows instructions with greater fidelity. It thinks more diligently, putting in more effort into multi-step planning and tool calls.

A more disciplined execution means less manual oversight and fewer retries across engineering workflows. For anyone running agentic loops in production, fewer retries is not a nice-to-have , it directly cuts compute cost and latency.

Google also demonstrated some striking real-world use cases for 3.7 Flash:

  • Generating a fully playable 3D game from a text prompt, with dynamic characters and textures created in real-time using Nano Banana.
  • Building interactive landing pages with parallax components in a single shot, using 3.7 Flash to orchestrate sub-agents.
  • Training a robotics model using multimodal understanding in a 3-agent graph loop.
  • Transforming static PDFs into interactive data stories with live charts.

Half the price, available now

3.7 Flash is available through the end of the year at an introductory price of $0.75/1M input tokens and $3.75/1M output tokens. To put that in context, 3.6 Flash was priced at $1.50 per 1M input tokens and $7.50 per 1M output tokens. That is a 50% cut across the board , while the model is meaningfully stronger on every benchmark Google published.

The introductory pricing does have a shelf life. Starting January 1, 2027, pricing will move to $1.50/1M input tokens and $7.50/1M output tokens , which is exactly what 3.6 Flash costs today. So the window to build and scale on the cheaper rate is real, but finite.

Where to use it

3.7 Flash is available across Google's full developer stack right now:

  • Developers: Explore agent-first workflows in Google Antigravity or start building today in the Gemini API via Google AI Studio and Android Studio.
  • Enterprises: Access via the Gemini Enterprise Agent Platform and the Gemini Enterprise app.
  • Individuals: Available via Spark, the 24/7 personal agent in the Gemini app for Google AI Pro and Ultra subscribers in over 160 countries.

Gemini Spark , Google's persistent personal agent that runs 24/7 and takes action on your behalf , is now powered by 3.7 Flash, making it more efficient for knowledge work with improved tool use for Google Workspace apps, delivering improved accuracy and output quality for complex, multi-skill workflows.

The Flash cadence is becoming a strategy

It is worth stepping back to look at what Google is doing with the Flash line. Google updates its Flash models at a rapid pace, with the Flash family evolving on a roughly bi-monthly cadence for lightweight, high-speed models. Each release is not a major architectural overhaul , it is a focused, targeted improvement shaped by real developer feedback.

This approach is working. The jump from 3.6 to 3.7 Flash on DeepSWE (49% to 65.3%) is larger than the jump from 3.5 to 3.6 Flash on the same benchmark. Google is compounding gains faster than the version numbers suggest. For teams building agents at scale, the practical implication is clear: the Flash tier is no longer a budget compromise. It is increasingly the right default choice for production agentic workloads , and at the introductory price, the economics are hard to argue with.

Safety

Gemini 3.7 Flash ships with updated safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense, while enabling beneficial use cases, in accordance with Google's approach to bioresilience and its cyber program.

Comments

avatar