Gemini 3.5 Flash API Pricing: Fast Multimodal Intelligence at Scale
Comprehensive Gemini 3.5 Flash API pricing analysis ($1.50/M input, $9.00/M output), 1M context caching, multimodal audio/video processing, and Google AI Studio SLAs.
How much does Gemini 3.5 Flash cost per million tokens?
Gemini 3.5 Flash costs $1.50 per million input tokens and $9.00 per million output tokens ($3.375/M blended at 3:1). Delivers low-latency multimodal processing across text, audio, images, and video with up to 1M context. Verified 2026-09-08.
How much does Gemini 3.5 Flash cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $0.6000 |
| Medium | 1,000 | 500 | $6.0000 |
| Long | 4,000 | 2,000 | $24.0000 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.
Volume ladder, output-cost sensitivity, and migration ledger for Gemini 3.5 Flash
1. Fixed-shape monthly cost ladder
| Workload | Input / output tokens | 100K requests | 1M requests | 10M requests |
|---|---|---|---|---|
| Short chat | 500 / 150 | $210.00 | $2100.00 | $21000.00 |
| Code review | 4,000 / 800 | $1320.00 | $13200.00 | $132000.00 |
| Document summary | 16,000 / 2,000 | $4200.00 | $42000.00 | $420000.00 |
Cost = requests × (input tokens × $1.50/M + output tokens × $9.00/M) ÷ 1,000,000, at three fixed text-token workload shapes. No volume discount, cache, or batch rate is inferred — Gemini 3.5 Flash's registry entry does not price those tiers.
2. Output-token cost-sensitivity band
| Workload | Input cost (1 request) | Output cost (1 request) | Output share of spend |
|---|---|---|---|
| Short chat | $0.0008 | $0.0014 | 64.3% |
| Code review | $0.0060 | $0.0072 | 54.5% |
| Document summary | $0.02 | $0.02 | 42.9% |
Output share = output-token spend ÷ (input-token spend + output-token spend) for a single request at each shape. Across these three shapes, Gemini 3.5 Flash's output share spans 42.9% to 64.3% — a 21.4%-point swing driven entirely by workload shape, not by any change in the $1.50/$9.00 per-1M rates.
3. Successor rate delta and lifecycle ledger
| Rate | Gemini 3.5 Flash | Gemini 3.6 Flash | Delta ($ / %) |
|---|---|---|---|
| Input $/M | $1.50 | $1.50 | $0.0000 (0.0%) |
| Output $/M | $9.00 | $7.50 | $-1.5000 (-16.7%) |
Delta = successor rate − Gemini 3.5 Flash rate, both taken from the current dated pricing registry. Cache, batch, and context-tier pricing are not substituted across models; each figure is model-specific or marked Unavailable.
| Field | Recorded value | Note |
|---|---|---|
| Lifecycle status | legacy | Verified 2026-08-14 |
| Deprecation announced | Unavailable | No inference beyond the dated record |
| Shutdown date | Unavailable | Null/unavailable is not a promise of indefinite availability |
| Successor | Gemini 3.6 Flash | gemini-3-6-flash priced in this registry |
| Context window | Unavailable | Unavailable — no model-specific spec sourced |
Price verified 2026-08-14; lifecycle verified 2026-08-14. Luna is the data owner. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated price source · lifecycle source · Test Gemini 3.5 Flash in All AI Ask.
gemini-3-5-flashGemini 3.5 Flash API Pricing: Fast Multimodal Intelligence at Scale
Gemini 3.5 Flash costs $1.50 per million input tokens and $9.00 per million output tokens ($3.375/M blended at 3:1). Delivers low-latency multimodal processing across text, audio, images, and video with up to 1M context. Verified 2026-09-08.
Gemini 3.5 Flash provides native multimodal token handling across text, audio, and visual inputs.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | High-resolution image analysis (1,500 tokens in, 200 out): $0.004050 per image |
| Scenario 2 | 1-minute audio recording transcription & summary (4K in, 500 out): $0.010500 per minute |
| Scenario 3 | 5-minute video keyframe extraction (20K in, 1K out): $0.039000 per video clip |
| Scenario 4 | PDF technical manual analysis (30K in, 2K out): $0.063000 per document |
| Scenario 5 | Customer video KYC verification (15K in, 500 out): $0.027000 per verification |
| Scenario 6 | Monthly multimodal processing tier (50M blended tokens): $168.75 infrastructure budget |
Google context caching slashes repetitive input processing costs on large media collections.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Cached video tutorial corpus (100K cached tokens, 2K query): 73% input cost reduction |
| Scenario 2 | Large product catalog cache (200K cached tokens): $0.07500 vs $0.30000 per search query |
| Scenario 3 | Multi-turn enterprise knowledge base Q&A: 70% net input savings across query bursts |
| Scenario 4 | Hourly storage fee ($4.50/M/hr) fully amortized after only 4 queries per hour |
| Scenario 5 | Latency reduction: cached queries bypass prompt processing stage, cutting TTFT by 50% |
| Scenario 6 | Enables interactive conversational exploration of heavy video and document archives |
Migrating to Gemini 3.6 or 3.7 Flash unlocks next-generation speed and a 95% token cost reduction.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Gemini 3.6/3.7 Flash pricing ($0.075/M in, $0.30/M out): 95% cheaper than 3.5 Flash |
| Scenario 2 | Gemini 3.7 Flash introduces native multimodal thinking tokens with superior reasoning |
| Scenario 3 | Migrating 50M tokens/mo saves $162.19/mo ($6.56 vs $168.75) with major latency gains |
| Scenario 4 | API compatibility: seamless migration on Google AI Studio and Vertex AI SDKs |
| Scenario 5 | Zero breaking changes in multimodal response formats or schema parameters |
| Scenario 6 | Strongly recommended action: immediate migration to Gemini 3.6 or 3.7 Flash |
Gemini 3.5 Flash pricing evidence
Source-backed dated rate shape: input=$1.500000/M; output=$9.000000/M. Historical Gemini 3.5 Flash rate and migration arithmetic only; 3.6/3.7 current economics, Google policy, and broad comparisons retain their owners. Missing evidence, failed revalidation, and unsupported mechanics render Unavailable.
Module 1 of 3: Historical audio/video-shaped bill ladder
Novel contribution boundary: Owns the dated 3.5 Flash token arithmetic for fixed multimodal-shaped inputs; audio/video unit pricing is not inferred. Formula / deterministic rule: datedSpend = requests × (inputTokens × inputRate + outputTokens × outputRate) / 1,000
| Scenario / field ID | Exact model, provider, and fixed inputs | Result | State |
|---|---|---|---|
batch83-gemini-3-5-flash-m1-r110K audio-transcript-shaped passes | model=gemini-3.5-flash; provider=google; registryRates=input=$1.500000/M; output=$9.000000/M; comparisonModel=gemini-3.6-flash; fixedInputs=requests=10,000; inputTokens=4,000; outputTokens=1,000; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicit | Unavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14. | FAIL CLOSED — no revalidated observation or provider guarantee |
batch83-gemini-3-5-flash-m1-r25K video-summary-shaped passes | model=gemini-3.5-flash; provider=google; registryRates=input=$1.500000/M; output=$9.000000/M; comparisonModel=gemini-3.6-flash; fixedInputs=requests=5,000; inputTokens=12,000; outputTokens=2,000; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicit | Unavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14. | FAIL CLOSED — no revalidated observation or provider guarantee |
batch83-gemini-3-5-flash-m1-r325K mixed summarization passes | model=gemini-3.5-flash; provider=google; registryRates=input=$1.500000/M; output=$9.000000/M; comparisonModel=gemini-3.6-flash; fixedInputs=requests=25,000; inputTokens=2,000; outputTokens=600; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicit | Unavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14. | FAIL CLOSED — no revalidated observation or provider guarantee |
First-party registry provenance: https://ai.google.dev/gemini-api/docs/pricing; registry owner=google; verifiedAt=2026-08-14; freshness gate=2026-09-14. Revalidate before any current-price claim.
Module 2 of 3: 3.5→3.6 accepted-result crossover
Novel contribution boundary: Owns dated migration arithmetic with explicit retry input; accepted-result quality and current-price status are not sourced. Formula / deterministic rule: requiredAcceptedRate = 3.6 Flash attempt-plus-retry cost / dated 3.5 attempt cost; observed acceptance = Unavailable
| Scenario / field ID | Exact model, provider, and fixed inputs | Result | State |
|---|---|---|---|
batch83-gemini-3-5-flash-m2-r1Audio-shaped summary; 10% retry share; 500 retry-input tokens | model=gemini-3.5-flash; provider=google; registryRates=input=$1.500000/M; output=$9.000000/M; comparisonModel=gemini-3.6-flash; fixedInputs=requests=1; inputTokens=4,000; outputTokens=1,000; retryInputTokens=500; retryShare=10%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicit | Unavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14. | FAIL CLOSED — no revalidated observation or provider guarantee |
batch83-gemini-3-5-flash-m2-r2Video-shaped summary; 20% retry share; 1,000 retry-input tokens | model=gemini-3.5-flash; provider=google; registryRates=input=$1.500000/M; output=$9.000000/M; comparisonModel=gemini-3.6-flash; fixedInputs=requests=1; inputTokens=12,000; outputTokens=2,000; retryInputTokens=1,000; retryShare=20%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicit | Unavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14. | FAIL CLOSED — no revalidated observation or provider guarantee |
batch83-gemini-3-5-flash-m2-r3Long mixed summary; 30% retry share; 2,000 retry-input tokens | model=gemini-3.5-flash; provider=google; registryRates=input=$1.500000/M; output=$9.000000/M; comparisonModel=gemini-3.6-flash; fixedInputs=requests=1; inputTokens=24,000; outputTokens=4,000; retryInputTokens=2,000; retryShare=30%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicit | Unavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14. | FAIL CLOSED — no revalidated observation or provider guarantee |
First-party registry provenance: https://ai.google.dev/gemini-api/docs/pricing; registry owner=google; verifiedAt=2026-08-14; freshness gate=2026-09-14. Revalidate before any current-price claim.
Module 3 of 3: Dated cache, endpoint, and lifecycle evidence panel
Novel contribution boundary: Owns dated 3.5 registry evidence only; cache, endpoint, and lifecycle facts are not fabricated. Formula / deterministic rule: evidence = dated exact-model registry field when present; missing field = Unavailable
| Scenario / field ID | Exact model, provider, and fixed inputs | Result | State |
|---|---|---|---|
batch83-gemini-3-5-flash-m3-r1Dated prompt-cache treatment | model=gemini-3.5-flash; provider=google; registryRates=input=$1.500000/M; output=$9.000000/M; comparisonModel=gemini-3.6-flash; fixedInputs=evidenceField=cache; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicit | Unavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14. | FAIL CLOSED — no revalidated observation or provider guarantee |
batch83-gemini-3-5-flash-m3-r2Dated endpoint or batch treatment | model=gemini-3.5-flash; provider=google; registryRates=input=$1.500000/M; output=$9.000000/M; comparisonModel=gemini-3.6-flash; fixedInputs=evidenceField=batch; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicit | Unavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14. | FAIL CLOSED — no revalidated observation or provider guarantee |
batch83-gemini-3-5-flash-m3-r3Dated lifecycle or retirement field | model=gemini-3.5-flash; provider=google; registryRates=input=$1.500000/M; output=$9.000000/M; comparisonModel=gemini-3.6-flash; fixedInputs=evidenceField=context; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicit | Unavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14. | FAIL CLOSED — no revalidated observation or provider guarantee |
First-party registry provenance: https://ai.google.dev/gemini-api/docs/pricing; registry owner=google; verifiedAt=2026-08-14; freshness gate=2026-09-14. Revalidate before any current-price claim.
How fast is Gemini 3.5 Flash?
How much does Gemini 3.5 Flash cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.34 |
| 1,000,000 | $3.38 |
| 10,000,000 | $33.75 |
| 100,000,000 | $337.50 |
How does Gemini 3.5 Flash compare with other models?
What should you explore next for Gemini 3.5 Flash?
What are common questions about Gemini 3.5 Flash?
Is Gemini 3.5 Flash cheaper than GPT-5?
Gemini 3.5 Flash costs $3.38/M blended tokens, GPT-5 costs $3.44/M — Gemini 3.5 Flash is cheaper.
How much does 1 million tokens cost with Gemini 3.5 Flash?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $3.38. Pure input costs $1.50/M; pure output costs $9.00/M.
What does Gemini 3.5 Flash cost at high volume?
At 100 million blended tokens a month, Gemini 3.5 Flash costs approximately $337.50. See the cost-at-scale table below for other volumes.
