Google Cuts Gemini Flash Price in Half Just Three Weeks After Launch

LLMsDeveloper Tools
Illustration generated by AI: Editorial image for Google Cuts Gemini Flash Price in Half Just Three Weeks After Launch

The Core · TL;DR

  • Gemini 3.7 Flash launches three weeks after Gemini 3.6 Flash at half the launch price: $0.75/$3.75 per million input/output tokens.
  • That introductory pricing expires December 31, 2026, rising to $1.50/$7.50 per million tokens in January 2027.
  • Coding and enterprise benchmarks jump sharply: DeepSWE rises from 49.0% to 65.3%, AutomationBench from 17.0% to 30.4%.
  • Google says the model beats Claude Sonnet 5 on its internal coding benchmarks; it's an algorithmic refinement of 3.6 Flash, not a new pretraining run.

Google has released Gemini 3.7 Flash barely three weeks after shipping Gemini 3.6 Flash, and the new model arrives at half the launch price of its predecessor. The rapid cadence signals just how aggressively Google is iterating within its "Flash" tier, the fast, low-cost branch of the Gemini family aimed at high-volume production workloads.

The pricing move is the headline detail. Gemini 3.7 Flash launches at $0.75 per million input tokens and $3.75 per million output tokens, half of what 3.6 Flash charged at its own debut.

That discount is introductory, however. According to Google's published rate card, the promotional pricing runs through December 31, 2026, after which costs double to $1.50 per million input tokens and $7.50 per million output tokens starting January 1, 2027. Buyers evaluating the model for long-term deployment should factor in that scheduled increase rather than assume the launch rate is permanent.

Coding gains drive the upgrade

Google frames 3.7 Flash as an algorithmic refinement of 3.6 Flash rather than a model trained from scratch on new pretraining data, but the benchmark jumps are substantial for a point release. On DeepSWE, a software-engineering benchmark, the new model scores 65.3% against 49.0% for its predecessor. On FrontierCode it climbs to 43.6% from 34.4%.

Enterprise-oriented evaluations show similar movement. AutomationBench, which measures multi-step workflow execution, rises to 30.4% from 17.0%, while GDP.pdf, a test of expert-level document comprehension, improves to 34.0% from 22.0%. Google also reports that 3.7 Flash outperforms Anthropic's Claude Sonnet 5 on its internal coding benchmarks, though that comparison comes from Google's own testing rather than independent verification.

The model keeps the same 1-million-token context window as its predecessor and can return up to 64K output tokens, with a knowledge cutoff of March 2026. It handles text, image, audio, and video inputs, and is available through the Gemini API, AI Studio, Antigravity, Android Studio, and Gemini Enterprise.

Why the pace matters

Three weeks between Flash releases is an unusually tight cycle even by frontier-lab standards, and it puts pressure on rivals to match both the performance gains and the price cuts. For developers building on Gemini, the practical takeaway is twofold: the coding and agentic-workflow improvements are real and benchmarked, but the attractive launch pricing has a firm expiration date built into Google's own documentation.

WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram