Gemini 3.5 Flash-Lite
Best for low-latency and high throughput agentic tasks
Best for low-latency and high throughput agentic tasks
Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second according to the Artificial Analysis Index
Handles tasks of varying complexity like coding, UI generation and translation with high quality
Delivers improved reasoning and output quality, allowing users to select the level of thinking they want to use.
Tackles tasks that require high throughput.
Uses search grounding and instruction following and built-in computer use.
A stronger price-to-performance ratio compared to 3.1 Flash-Lite.
Explore what you can do with Gemini 3.5 Flash-Lite
3.5 Flash-Lite executes high volume tasks at a lower latency than 3.5 Flash
Working alongside 3.6 Flash as the master agent, 3.5 Flash-Lite instantly generates 25 unique, ready-to-explore web design concepts
3.5 Flash-Lite can scale receipt translation and summarization with its multimodal understanding
3.5 Flash-Lite builds a game by instantly generating and iterating through multiple options
3.5 Flash-Lite enables efficient scaling for agentic systems. Across thinking levels, the model significantly outperforms 3.1 Flash-Lite. In fact, on many agentic and coding evals, 3.5 Flash-Lite even outperforms 3 Flash.
| Benchmark | Notes | Gemini 3.5 Flash-Lite | Gemini 3.1 Flash-Lite | GPT-5.4 mini | Claude Haiku 4.5 |
|---|---|---|---|---|---|
| Input price $/1M tokens, no caching | $0.30 | $0.25 | $0.75 | $1.00 | |
| Output price $/1M tokens | $2.50 | $1.50 | $4.50 | $5.00 | |
| SWE-Bench Pro (Public) Diverse agentic coding tasks | 54.2% | 38.3% | 54.4% | 39.5% | |
| Terminal-bench 2.1 Agentic terminal coding | Terminus-2 harness | 54.0% | 31.0% | 59.2% | 44.2% |
| MLE-Bench Machine Learning Engineering | 39.2% | 22.0% | — | — | |
| GDPVal-AA v2 Knowledge work | Elo | 1140 | 642 | 1171 | 907 |
| OSWorld-Verified Agentic computer use | 74.0% | 54.3% | 72.1% | 50.7% | |
| CharXiv Reasoning Information synthesis from complex charts | No tools | 74.5% | 73.2% | 80.3% | 61.7% |
| With tools | 76.5% | 75.6% | — | — | |
| GDM-MRCR v2 (8-needle) Long context performance | 128k (average) | 72.2% | 60.1% | 42.7% | 35.3% |
| 1M (pointwise) | 21.3% | 12.3% | — | — |