
Gemini 3.7 Flash is a low-cost coding and agent model from Google for developers who want fast, cheap model access without paying premium rates. Released on 13 August 2026, it is the third Flash-tier iteration in three weeks and is built on Gemini 3.6 Flash with a reworked reasoning core tuned for software engineering and autonomous workflows.
Core Features
- A one-million-token context window with a 64,000-token output limit, matching the 3.6 Flash envelope.
- Three configurable thinking levels (low, medium, high) so you trade latency for accuracy per request.
- Native tool use: function calling, code execution, search and Maps grounding, URL context, and structured outputs.
- Strong coding and agent scores: 65.3 percent on DeepSWE v1.1 and 43.6 percent on FrontierCode 1.1, up sharply from 3.6 Flash.
- Multimodal input across text, images, audio, video, and PDF, plus computer use in preview.
Use Cases
- Building and debugging software agents that plan, call tools, and verify results across long contexts.
- Powering real-time coding assistants and in-IDE completions where speed and cost matter more than max reasoning.
- Running high-volume document and PDF comprehension pipelines on a budget.
Pricing
Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens through 31 December 2026. From 1 January 2027 the rate doubles to $1.50 and $7.50. A free tier is available inside the Gemini app via Spark for AI Pro and Ultra subscribers, and the API runs in Google AI Studio and Vertex. Context caching is $0.075 per million tokens.
Our Take
Best for developers and startups who need Claude Haiku-class speed at roughly half the price and can live with Google's new thinking-level API. The trade-off is that it lags GPT-5.6 and DeepSeek V4-Flash on raw cost per token, and it is not the model to pick for heavy open-ended reasoning. Browse more options in our AI model directory.




