
Google's Gemini 3.6 Flash is now rolling out in GitHub Copilot. It is the latest in Google's Flash line -- models tuned for the sweet spot between raw intelligence and the speed and cost that production workloads actually demand. The pitch this time: better code, fewer tokens burned, and a lower price tag than the model it replaces.
What changed from 3.5 Flash
3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%), and shows significant improvement in ML Research, as seen in MLE Bench (63.9% vs. 49.7%). Those are meaningful jumps -- DeepSWE measures an agent's ability to resolve real software engineering tasks end-to-end, and MLE Bench tests autonomous ML research capabilities.
The gains aren't limited to coding. It has improved computer use capabilities as seen in OSWorld-Verified (83.0% vs. 78.4%), and it outperforms 3.5 Flash in knowledge work, as shown by benchmarks like GDPval-AA v2 (1421 vs. 1349). GDPval-AA is an agentic benchmark that scores models on real-world economically valuable tasks using an Elo rating system, similar to chess rankings.
On OSWorld-Verified, a computer-use benchmark, Gemini 3.6 Flash actually posts the best score in the entire comparison at 83.0%, ahead of every other model listed including GPT 5.6 Luna and Grok 4.5. Long-context retrieval is another standout: on GDM-MRCR v2 at 1M tokens, Gemini 3.6 Flash scores 54.0% while Gemini 3.5 Flash and Gemini 3.1 Pro don't clear 27%.
Don't miss what's next in AI
Join 300,000+ engineers and researchers who get the signal, not the noise.
- Full access to in-depth AI research breakdowns
- Be the first to know what's trending before it hits mainstream
- Daily curated papers, repos, and industry moves

