Gemini 3.7 Flash β€” Google's best-value frontier model for coding and agents

Gemini 3.7 Flash is Google DeepMind's workhorse model, tuned for long-running coding agents, document-heavy knowledge work, and native understanding of text, image, audio, and video.

Gemini 3.7 Flash β€” Google's best-value frontier model for coding and agents
In / Out price
$0.75 / $3.75 per 1M
Context
1M
Released
Aug 13, 2026
Knowledge cutoff
March 2026

Frontier coding at workhorse prices

A refinement of 3.6 Flash that closed most of the gap to the premium tier, at roughly a third of the blended cost.

Long-running coding agents

Long-running coding agents

DeepSWE v1.1 jumped from 49.0% to 65.3% in three weeks. FrontierCode 1.1 Main went from 34.4% to 43.6%. The model holds its plan across multi-file refactors instead of losing the thread halfway.

Front-end that actually renders

Front-end that actually renders

1588 Elo on WebDev Arena, up 50 points from 3.6 Flash and ahead of Claude Sonnet 5 at 1541. Feed it a screenshot or a design system and it builds against it.

Reasoning you can meter

Reasoning you can meter

Three thinking levels β€” low, medium, high β€” so you decide per call whether you're buying quality or latency. Batch jobs run cheap; the hard calls get the full budget.

Documents, not just text

Documents, not just text

34.0% on GDP.pdf against 22.0% for the previous release, plus 90.7% on Harvey's legal benchmark. PDFs, contracts, and financial filings go in; structured data comes out.

Experience The Brilliance

Gemini 3.7 Flash turns the work that used to take a whole afternoon into something you kick off and come back to.

Ship a Feature, Not a Snippet

Ship a Feature, Not a Snippet

Most models help you write a function. This one takes a ticket, reads the codebase, opens the files it needs, and works through the change. AutomationBench, which measures exactly this kind of multi-step business workflow, went from 17.0% to 30.4% between 3.6 and 3.7 Flash. Claude Sonnet 5 sits at 10.7% on the same test.

Read the Screenshot, Build the Page

Read the Screenshot, Build the Page

Point it at a Figma frame, a competitor's landing page, or a whiteboard photo, and ask for the markup. Video works too: 85.4% on LVBench means it follows what happens across a long clip rather than sampling a few frames and guessing.

Pay for the Thinking You Need

Pay for the Thinking You Need

The thinking configuration is a dial, not a switch. Set it low for classification and extraction where the answer is either right or wrong. Set it high when you're handing over a debugging session and want the model to sit with the problem.

One model, multiple capabilities

Everything below runs off the same endpoint. There's no separate model for vision, no separate model for tools, and nothing to request access to except computer use.

1M-token context window

The model takes 1,048,576 tokens in and returns up to 65,536, so a mid-sized repository fits in a single request.

Four input modalities

Text, images, audio, and video all go into the same prompt, and PDFs parse natively without a separate extraction step.

Three thinking levels

You pick low, medium, or high per call, trading quality against cost and latency wherever it makes sense.

Function calling

Hand it your schema and it selects the right tool and fills the arguments, which is enough to wire it into APIs you already run.

Code execution

It writes code, runs it in a sandbox, and reads the result before answering, catching arithmetic and logic slips before they reach you.

Computer use

It drives a browser and a desktop directly, clicking through the interfaces that never shipped an API.

File search

Upload a corpus and the model retrieves against it, citing the source document rather than paraphrasing from memory.

Search grounding

Live Google Search results land in the response, so answers don't stop at the March 2026 cutoff.

Maps grounding

Location questions resolve against Google Maps data instead of whatever the training run happened to absorb.

Have questions?
We have answers!

Google DeepMind's newest workhorse model, released August 13, 2026. It's built on Gemini 3.6 Flash with algorithmic improvements to the reasoning foundation rather than a new pretraining run. Google positions it for coding, agentic workflows, and enterprise knowledge work. It takes text, images, audio, video, and PDFs, and returns text.

Through December 31, 2026, the API rate is $0.75 per 1M input tokens and $3.75 per 1M output. That's half what 3.6 Flash launched at. From January 1, 2027 it moves to $1.50 / $7.50. On a typical 80/20 input-output mix that works out to about $1.35 per million now, against roughly $3.60 for Claude Sonnet 5 and $4.00 for GPT-5.6 Terra. Through Imagine Computer you get it as part of the plan at $9/month billed yearly, or $13/month billed monthly.

No. Google ships it through the Gemini API, AI Studio, Antigravity, Android Studio, and the Gemini Enterprise Agent Platform. Google has never released weights for a Gemini model and hasn't for this one, so self-hosting and air-gapped deployment aren't options. Consumers reach it through Gemini Spark on AI Pro and Ultra plans in over 160 countries, though the rollout excludes the European Economic Area, the UK, Switzerland, and Nigeria.

It wins on web development (1588 Elo on WebDev Arena, the highest of the current frontier set) and on business automation (30.4% on AutomationBench versus 23.6% for GPT-5.6 Terra). At the high thinking setting it scores 56 on the Artificial Analysis Intelligence Index, effectively level with Claude Sonnet 5 at 55 and GPT-5.6 Terra and Muse Spark 1.2 at 57. GPT-5.6 Terra still leads on terminal and computer-use agents: 69.6% on DeepSWE, 87.4% on Terminal-bench 2.1, and 50.2% on OSWorld-2.0. And on GDPval-AA v2, which measures general knowledge work, Gemini 3.7 Flash comes in last of the four at 1525, behind Muse Spark 1.2 (1628), Claude Sonnet 5 (1598), and GPT-5.6 Terra (1578).

Two things worth putting in your system prompt. Retrieval degrades at the top of the context window β€” 97.0% at 128k tokens drops to 62.5% at the full 1M, so chunk long inputs rather than dumping them. And chart reasoning regressed slightly from 3.6 Flash (84.5% versus 85.2% on CharXiv), so verify anything downstream of a figure. Google's model card also flags occasional slowness and timeouts under load, plus the usual hallucination risk. Knowledge cutoff is March 2026, though some domains only extend to January 2025.

Try Gemini Flash Models in Imagine Computer

We've added Gemini Flash Models to Imagine Computer's [AI Chat](imagine.art/imagine-computer/ai-chat), so Google's newest coding and agent model runs directly in your workspace alongside everything else you use.