Google's Gemini 3.6 Flash Hits GitHub Copilot Faster and 17% Cheaper

Google's Gemini 3.6 Flash lands in GitHub Copilot with better coding benchmarks, 17% fewer output tokens, and a lower price than its predecessor

·
·
Google's Gemini 3.6 Flash Hits GitHub Copilot Faster and 17% Cheaper
AuthorGitHub
Read2 min
  • Gemini 3.6 Flash is now rolling out in GitHub Copilot for Pro, Pro+, Max, Business, and Enterprise users.
  • Coding benchmark scores jump significantly: DeepSWE goes from 37% to 49%, MLE Bench from 49.7% to 63.9% vs. 3.5 Flash.
  • 17% fewer output tokens than 3.5 Flash on average, up to 65% fewer on some benchmarks -- directly cutting agentic task costs.
  • Priced at $1.50/1M input and $7.50/1M output tokens -- cheaper than 3.5 Flash's $9 output price despite better performance.
  • Features configurable reasoning effort, parallel tool use, 1M token context window, and built-in computer use.
  • Enterprise/Business admins must enable the Gemini 3.6 Flash Preview policy in Copilot settings before users can access it.

Google's Gemini 3.6 Flash is now rolling out in GitHub Copilot. It is the latest in Google's Flash line -- models tuned for the sweet spot between raw intelligence and the speed and cost that production workloads actually demand. The pitch this time: better code, fewer tokens burned, and a lower price tag than the model it replaces.

What changed from 3.5 Flash

3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%), and shows significant improvement in ML Research, as seen in MLE Bench (63.9% vs. 49.7%). Those are meaningful jumps -- DeepSWE measures an agent's ability to resolve real software engineering tasks end-to-end, and MLE Bench tests autonomous ML research capabilities.

The gains aren't limited to coding. It has improved computer use capabilities as seen in OSWorld-Verified (83.0% vs. 78.4%), and it outperforms 3.5 Flash in knowledge work, as shown by benchmarks like GDPval-AA v2 (1421 vs. 1349). GDPval-AA is an agentic benchmark that scores models on real-world economically valuable tasks using an Elo rating system, similar to chess rankings.

On OSWorld-Verified, a computer-use benchmark, Gemini 3.6 Flash actually posts the best score in the entire comparison at 83.0%, ahead of every other model listed including GPT 5.6 Luna and Grok 4.5. Long-context retrieval is another standout: on GDM-MRCR v2 at 1M tokens, Gemini 3.6 Flash scores 54.0% while Gemini 3.5 Flash and Gemini 3.1 Pro don't clear 27%.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves