Google's Gemini 3.6 Flash Cuts Output Costs 17% While Beating Coding Benchmarks
Google's July Gemini Drops brings faster Flash models, system-wide macOS voice control, Spark's global rollout, and personalized image generation for all U.S. users.

- Gemini 3.6 Flash replaces 3.5 Flash as the default workhorse model: 17% fewer output tokens, better coding benchmarks (DeepSWE 49% vs 37%), priced at $1.50/$7.50 per 1M tokens.
- Gemini 3.5 Flash-Lite launches at $0.30/$2.50 per 1M tokens, running at 350 tokens/sec and outperforming the older 3 Flash on agentic and coding evals.
- Gemini for macOS gets system-wide voice control via Fn key hold, with intelligent dictation, screen-aware rewriting, and image editing across any open window.
- Gemini Spark expands to 160+ new countries for paid subscribers and gains Chrome integration to handle web errands using saved accounts.
- Personalized image generation via Nano Banana is now free for all eligible U.S. users, with a one-time digital avatar setup for selfie-free image creation.
- New app integrations added: Dropbox, Viator, and Zillow; Gemini 3.5 Pro still absent, confirmed still in partner testing.
Google's July Gemini Drops covers a lot of ground: two new Flash models with real benchmark improvements, a macOS voice feature that works across every open app, a major geographic expansion for the Gemini Spark agent, new third-party integrations, and a digital avatar system for personalized image generation.
Two new Flash models, cheaper and faster
The headline update is Gemini 3.6 Flash, which replaces 3.5 Flash as Google's everyday workhorse. It promises stronger performance in coding, knowledge work, and multimodal tasks while cutting token usage by up to 17%. Concretely, 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens, down from 3.5 Flash's $9.00 output rate.
The efficiency gains go beyond pricing. Gemini 3.6 Flash keeps the 1 million-token context window, advances the knowledge cutoff to March 2026, and scores higher on coding, long-context, and computer-use benchmarks. On coding tasks it delivers fewer unwanted edits and shorter execution loops, with benchmark scores of 49% on DeepSWE (up from 37%) and 63.9% on MLE Bench (up from 49.7%).
The second model, Gemini 3.5 Flash-Lite, targets high-throughput, low-latency workloads such as agentic search and document processing. At $0.30 per million input tokens and $2.50 per million output tokens, it's the cheapest option in the lineup, yet it outperforms 3 Flash on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%).