Gemini 3.6 Flash
Best for token efficiency in coding, knowledge work, and multimodal tasks
Best for token efficiency in coding, knowledge work, and multimodal tasks
Our workhorse model that reduces output token usage by 17% compared to 3.5 Flash, according to Artificial Analysis Index.
Get advanced reasoning at Flash-level latency and scale.
Get better quality in coding, knowledge work, and multimodal tasks whilst reducing token usage.
Speed and scale don’t have to come at the cost of intelligence.
Deep reasoning across long horizons and iterative coding tasks.
Multimodal understanding across text, audio, images, code, and video.
Here are a few ways you can use 3.6 Flash’s improved coding and knowledge work capabilities.
3.6 Flash shows better token efficiency and reduced verbosity than 3.5 Flash in multi-step workflows, as seen in the OSWorld verified task
3.6 Flash, using Managed Agents on AIS, can help parse through and analyze financial data and transcripts more efficiently and accurately than 3.5 Flash
3.6 Flash executes code migrations, using multi-agent orchestration on AGY, with lower latency and higher quality than 3.5 Flash
3.6 Flash helps develop a photographic texture extractor for 3D workflows, using canvas
3.6 Flash is more token efficient than 3.5 Flash and a step-up in coding and knowledge work.
| Benchmark | Notes | Gemini 3.6 Flash | Gemini 3.5 Flash | Gemini 3.1 Pro | GPT-5.6 Luna | Grok 4.5 | Claude Sonnet 5 |
|---|---|---|---|---|---|---|---|
| Input price $/1M tokens, no caching | $1.50 | $1.50 | $2.00 | $1.00 | $2.00 | $3.00 $2.00 (temp discount) | |
| Output price $/1M tokens | $7.50 | $9.00 | $12.00 | $6.00 | $6.00 | $15.00 $10.00 (temp discount) | |
| SWE-Bench Pro (Public) Diverse agentic coding tasks | 58.7% | 55.1% | 54.2% | 62.7% | 64.7% | 63.2% | |
| DeepSWE v1.1 Long-horizon software engineering | 49% | 37% | 12% | 67% | 54% | 54% | |
| Terminal-bench 2.1 Agentic terminal coding | Terminus-2 harness | 78.0% | 76.2% | 73.8% | 84.7% | 83.3% | 80.4% |
| MLE-Bench Machine Learning Engineering | 63.9% | 49.7% | 42.6% | 47.6% | 43.2% | 66.9% | |
| OSWorld-Verified Agentic computer use | 83.0% | 78.4% | 76.2% | 72.6% | — | 81.2% | |
| GDPVal-AA v2 Knowledge work | Elo | 1421 | 1349 | 965 | 1584 | 1535 | 1607 |
| CharXiv Reasoning Information synthesis from complex charts | No tools | 85.2% | 84.2% | 83.3% | 82.7% | 81.6% | 77.0% |
| With tools | 89.4% | 84.9% | 83.2% | — | — | 88.3% | |
| GDM-MRCR v2 (8-needle) Long context performance | 128k (average) | 91.8% | 77.3% | 84.9% | 74.8% | 81.4% | 71.6% |
| 1M (pointwise) | 54.0% | 26.6% | 26.3% | — | — | — |
For details on our evaluation methodology please see deepmind.google/models/evals-methodology/gemini-3-6-flash