Introducing our latest series of models combining frontier intelligence with action. Build more capable, intelligent agents efficiently.

Models

Completing everyday tasks, or solving your most challenging problems. Discover the right model for what you need

Performance

BenchmarkNotesGemini 3.6 FlashGemini 3.5 FlashGemini 3.1 ProGPT-5.6 LunaGrok 4.5Claude Sonnet 5
Input price $/1M tokens, no caching$1.50$1.50$2.00$1.00$2.00$3.00 $2.00 (temp discount)
Output price $/1M tokens$7.50$9.00$12.00$6.00$6.00$15.00 $10.00 (temp discount)
SWE-Bench Pro (Public) Diverse agentic coding tasks58.7%55.1%54.2%62.7%64.7%63.2%
DeepSWE v1.1 Long-horizon software engineering49%37%12%67%54%54%
Terminal-bench 2.1 Agentic terminal codingTerminus-2 harness78.0%76.2%73.8%84.7%83.3%80.4%
MLE-Bench Machine Learning Engineering63.9%49.7%42.6%47.6%43.2%66.9%
OSWorld-Verified Agentic computer use83.0%78.4%76.2%72.6%81.2%
GDPVal-AA v2 Knowledge workElo14211349965158415351607
CharXiv Reasoning Information synthesis from complex chartsNo tools85.2%84.2%83.3%82.7%81.6%77.0%
With tools89.4%84.9%83.2%88.3%
GDM-MRCR v2 (8-needle) Long context performance128k (average)91.8%77.3%84.9%74.8%81.4%71.6%
1M (pointwise)54.0%26.6%26.3%

For details on our evaluation methodology please see deepmind.google/models/evals-methodology/gemini-3-6-flash




Safety

Building with responsibility at the core

As we develop these new technologies, we recognize the responsibility it entails, and aim to prioritize safety and security in all our efforts.


Gemini Ecosystem


Try Gemini