Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second according to the Artificial Analysis Index



Performance

3.5 Flash-Lite enables efficient scaling for agentic systems. Across thinking levels, the model significantly outperforms 3.1 Flash-Lite. In fact, on many agentic and coding evals, 3.5 Flash-Lite even outperforms 3 Flash.

BenchmarkNotesGemini 3.5 Flash-LiteGemini 3.1 Flash-LiteGPT-5.4 miniClaude Haiku 4.5
Input price $/1M tokens, no caching$0.30$0.25$0.75$1.00
Output price $/1M tokens$2.50$1.50$4.50$5.00
SWE-Bench Pro (Public) Diverse agentic coding tasks54.2%38.3%54.4%39.5%
Terminal-bench 2.1 Agentic terminal codingTerminus-2 harness54.0%31.0%59.2%44.2%
MLE-Bench Machine Learning Engineering39.2%22.0%
GDPVal-AA v2 Knowledge workElo11406421171907
OSWorld-Verified Agentic computer use74.0%54.3%72.1%50.7%
CharXiv Reasoning Information synthesis from complex chartsNo tools74.5%73.2%80.3%61.7%
With tools76.5%75.6%
GDM-MRCR v2 (8-needle) Long context performance128k (average)72.2%60.1%42.7%35.3%
1M (pointwise)21.3%12.3%

Model information

Name
3.5 Flash-Lite
Status
General availability
Input
  • Text
  • Image
  • Video
  • Audio
  • PDF
Output
  • Text
Input tokens
1M
Output tokens
64k
Tool use
  • Function calling
  • Search as a tool
  • Computer use
Best for
  • High-volume, latency-sensitive reasoning tasks
Availability
  • Gemini App
  • Google AI Studio
  • Gemini Enterprise Agent Platform
  • Gemini API
Documentation
View developer docs
Model card
View model card