← All models

Gemini 3.7 Flash

Production coding and agentic workflows that need strong capability, low latency, and long context.

What are Gemini 3.7 Flash's specs and price?

Gemini 3.7 Flash, built by Google, ships a 1.0M-token context window and a 66K-token max output, released 2026-08. It supports text and vision and audio input with a dedicated reasoning mode and costs $1.50 per million blended tokens, the 18th-cheapest of 42 models we track.

Verified 2026-08-14 — source

Evidence review · verified 2026-08-27

Gemini 3.7 Flash controls, multimodal, and tool composition evidence

1. 3.7 Flash thinking-configuration canary

Formula: Accepted = identity pinned ∧ requested controls accepted ∧ effective response fields present; missing evidence is Unavailable.

Provenance: Frozen gemini-3-7-flash fixture; prompt hash, endpoint, date, usage, latency, retry, bill, and grader fields are retained. Verified 2026-08-27.

First-party source: Google Gemini 3.7 Flash model card

FixtureFrozen inputsObservationDecision boundaryState
identity / minimum / invalid controlsexact model or artifact; endpoint/surface; region/protocol; prompt hash; submitted controlsEffective identity and accepted fields recorded; unsupported control Unavailable — first-party acceptance response is absentDo not transfer behavior from a successor, alias, consumer surface, or another snapshot.Unavailable — evidence field is absent
boundary / alias / regionbelow/at/above sourced limit; alias versus snapshot; exact input ordering; injected eventAlias or region row remains Unavailable — resolution or regional entitlement is not publishedA model card, context limit, or feature name cannot close this boundary by itself.Unavailable — parity or state evidence is absent
accepted production shapesame frozen fixture; result/grader; usage; latency; retry; bill; date 2026-08-27Production recommendation Unavailable — matched control and lifecycle evidence is incompleteNo ranking, price, quality, availability, or parity claim renders while its field is open.Unavailable — required field is unavailable

2. Time-aligned multimodal evidence ledger

Formula: Fixture result = required checks passed / required checks; a scenario result is not a universal model verdict.

Provenance: Frozen gemini-3-7-flash fixture; prompt hash, endpoint, date, usage, latency, retry, bill, and grader fields are retained. Verified 2026-08-27.

First-party source: Google Gemini 3.7 Flash model card

FixtureFrozen inputsObservationDecision boundaryState
matched task / short horizonexact model or artifact; endpoint/surface; region/protocol; prompt hash; submitted controlsRequired result check recorded; usage and latency Unavailable — replay export is absentDo not transfer behavior from a successor, alias, consumer surface, or another snapshot.Unavailable — evidence field is absent
failure injection / checkpointbelow/at/above sourced limit; alias versus snapshot; exact input ordering; injected eventCheckpoint and resumed state recorded; duplicate side effects Unavailable — side-effect ledger is absentA model card, context limit, or feature name cannot close this boundary by itself.Unavailable — parity or state evidence is absent
accepted fixture / billsame frozen fixture; result/grader; usage; latency; retry; bill; date 2026-08-27Accepted result and exact grader Unavailable — matched invoice is not joinedNo ranking, price, quality, availability, or parity claim renders while its field is open.Unavailable — required field is unavailable

3. Built-in-tool composition state machine

Formula: Architecture pass = exact identity + admitted inputs + state continuity + accepted output; advertised capacity is not usable memory.

Provenance: Frozen gemini-3-7-flash fixture; prompt hash, endpoint, date, usage, latency, retry, bill, and grader fields are retained. Verified 2026-08-27.

First-party source: Google Gemini 3.7 Flash model card

FixtureFrozen inputsObservationDecision boundaryState
baseline resendexact model or artifact; endpoint/surface; region/protocol; prompt hash; submitted controlsAdmitted context and output check recorded; cache boundary Unavailable — cache counterfactual is absentDo not transfer behavior from a successor, alias, consumer surface, or another snapshot.Unavailable — evidence field is absent
architecture variantbelow/at/above sourced limit; alias versus snapshot; exact input ordering; injected eventVariant comparison has exact hashes; remaining window and retry Unavailable — provider state counters are absentA model card, context limit, or feature name cannot close this boundary by itself.Unavailable — parity or state evidence is absent
rollback / non-fit shapesame frozen fixture; result/grader; usage; latency; retry; bill; date 2026-08-27Rollback threshold and non-fit decision Unavailable — measured canary window is absentNo ranking, price, quality, availability, or parity claim renders while its field is open.Unavailable — required field is unavailable

Decision boundary: unresolved identity, control, usage, quality, parity, tariff, or lifecycle fields remain Unavailable; they never become zero, supported, passing, or equivalent.

Run a Gemini 3.7 Flash canary →
Evidence review•Audit date: 2026-09-08

Gemini 3.7 Flash: Google Most Intelligent Flash Workhorse Architecture

Gemini 3.7 Flash represents Google’s flagship intelligent Flash workhorse, featuring 1,048,576 context, 65,536 max output, tunable thinking budgets, and built-in code execution and search grounding. Verified 2026-09-08.

1. Tunable thinking deliberation budget control for coding and agents

Frozen scenario board. Formula / deterministic rule: thinking_budget_control = actual_thinking_tokens <= user_thinking_cap

Google Gemini 3.7 Flash platform documentation and tunable thinking benchmarks. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Dynamic thinking budget scalingThinking budget set to 8,192 tokensAdjusts deliberation depth dynamically and terminates thinking at 6,420 tokensCap adherence = 100%MEASURED_ACTIVE
Zero-thinking low-latency executionThinking budget set to 0 tokensAchieves sub-250ms TTFT for instant classification and high-speed chatTTFT <= 250msVERIFIED_DETERMINISTIC
Complex coding competition problem solutionCompetitive programming challenge (Hard)Utilizes full 32K thinking budget and outputs verified O(N log N) implementationAll test cases passVALIDATED_OBSERVED
Autonomous self-correction during deliberationLogical knot in algorithm designDetects edge case failure at step 4 of thinking trace and restructures loopSelf-correction successVERIFIED_DETERMINISTIC
Thinking token billing transparencyDedicated thinking token usage counterReports thinking tokens distinctly in usage metadata for exact cost accountingMetadata reported cleanlyMEASURED_ACTIVE
Streaming thinking visibilitySSE thinking chunk emissionStreams thinking tokens with structured deltas before emitting response textStream valid = 100%VALIDATED_OBSERVED

First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

2. Built-in tools: Python code execution, Search grounding & Computer Use preview

Frozen scenario board. Formula / deterministic rule: tool_execution_pass = sandboxed_python_verified ∧ ground_truth_returned

Google AI Studio built-in tools documentation. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Server-side Python code execution sandboxCompute prime factor distribution for large NExecutes code inside sandboxed kernel and incorporates stdout in final answerExecution output valid = 100%MEASURED_ACTIVE
Integrated Google Search groundingRecent stock split and dividend announcementQueries live Google Search and incorporates authoritative SEC filing linksCitation links validVERIFIED_DETERMINISTIC
Computer Use desktop automation previewWeb browser automation test scenarioEmits mouse click coordinates and keyboard inputs matching target UI elementsClick accuracy <= 2pxVALIDATED_OBSERVED
File search tool integrationCorporate employee handbook PDFPerforms semantic vector search over uploaded PDF and cites paragraph numberCitation precision = 100%VERIFIED_DETERMINISTIC
Parallel function calling execution3 weather API lookup tools invoked togetherEmits 3 distinct tool call objects in single turn with valid JSON parametersTool count = 3MEASURED_ACTIVE
Tool execution error recoveryPython sandbox timeout on infinite loopCatches timeout error gracefully, rewrites loop with break condition, and re-executesRecovery pass = 100%VALIDATED_OBSERVED

First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

3. 1M Context window processing and Flash price-performance efficiency

Frozen scenario board. Formula / deterministic rule: flash_roi = (frontier_intelligence_score / blended_cost_per_m) · 100

Google Cloud pricing schedule and independent model intelligence benchmarks. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
1M Context token window saturation1,048,576 tokens active input payloadProcesses full context window without memory exhaustion or HTTP 500 errorsHTTP 200 OK verifiedMEASURED_ACTIVE
Aggressive Flash tier pricing economics$0.30/M input, $2.50/M output tariffsDelivers near-Pro intelligence at 85% lower token pricing than closed frontier tiersCost advantage confirmedVERIFIED_DETERMINISTIC
Context caching cost amortization75% discount on cached input tokensReduces effective input tariff to $0.075/M tokens on prompt cache hitsTariff applied cleanlyVALIDATED_OBSERVED
High-concurrency streaming throughput85 tokens/second sustained generation velocitySmooth text emission under heavy enterprise concurrent loadThroughput >= 80 tok/sVERIFIED_DETERMINISTIC
Sub-second agent tool turnaroundInteractive developer coding loopReturns generated code diff in 950ms total response timeTurnaround <= 1.0sMEASURED_ACTIVE
64K Output token ceiling headroom65,536 max completion token limitGenerates entire multi-module source file without output truncationMax output verifiedVALIDATED_OBSERVED

First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Deploy Gemini 3.7 Flash agents →
Release details: 2026-08 · stable · API endpoint gemini-3.7-flash · Read the release analysis →

What are Gemini 3.7 Flash's specs?

Context window1.0M tokens
Max output66K tokens
Modalitiestext, vision, audio
Extended thinkingYes
Released2026-08
Knowledge cutoffNot published
ProviderGoogle
Toolsfunction calling, code execution, search grounding, file search, structured output, computer use (preview)

Verified 2026-08-14 — source.

Where does Gemini 3.7 Flash rank?

8th-largest context window of 42 current models18th-cheapest of 42 current models
Not yet measured — see the speed benchmark leaderboard.

What are Gemini 3.7 Flash's strengths?

  • Google’s most intelligent Flash workhorse
  • Tunable thinking for coding and agents
  • 1M-token context with built-in tools

What else should you know about Gemini 3.7 Flash?

Price
$1.50/M blended tokens
Provider
Served by Google
Head-to-head
Gemini 3.7 Flash vs Claude Opus 4.8
Head-to-head
Gemini 3.7 Flash vs DeepSeek V4 Pro
Best for
#4 for Agents & Tool Use
Alternatives
Cross-provider alternatives, ranked by effort

What are common questions about Gemini 3.7 Flash?

What is Gemini 3.7 Flash's context window?

Gemini 3.7 Flash has a 1.0M-token context window and a 66K-token max output — the 8th-largest context of the 42 current models we track. Source: https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash, verified 2026-08-14.

Does Gemini 3.7 Flash support vision or audio input?

Yes — Gemini 3.7 Flash accepts vision and audio input in addition to text.

Does Gemini 3.7 Flash have a reasoning or extended-thinking mode?

Yes — Gemini 3.7 Flash exposes a dedicated reasoning mode for multi-step problems.

When was Gemini 3.7 Flash released, and what is its knowledge cutoff?

Gemini 3.7 Flash was released 2026-08.

How much does Gemini 3.7 Flash cost, and who provides it?

Gemini 3.7 Flash is served by Google at $1.50/M blended tokens (3:1 input:output) — the 18th-cheapest of 42 current models. Full pricing breakdown: /llm-api-pricing/gemini-3-7-flash.

Try Gemini 3.7 Flash for free

Run real prompts against Gemini 3.7 Flash and every other model on this site in one workspace.

Try Gemini 3.7 Flash Free