← Back to all pricing

DeepSeek V4 Flash API Pricing: Ultra-Low-Cost High-Speed Intelligence

Comprehensive DeepSeek V4 Flash API pricing analysis ($0.44/M input, $1.32/M output), off-peak discounts, classification speed, and prompt caching breaks.

Full specs, context window and API limits →

How much does DeepSeek V4 Flash cost per million tokens?

DeepSeek V4 Flash costs $0.44 per million input tokens and $1.32 per million output tokens ($0.66/M blended at 3:1). A high-efficiency open-weights model designed for rapid classification, document summarization, and large-scale data structuring. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$0.44/M
Output
$1.32/M
Blended
$0.66/M
Provider
Verified 2026-08-14source

How much does DeepSeek V4 Flash cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 2.59× verbosity factor.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.2149
Medium1,000500$2.1494
Long4,0002,000$8.5976

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.

Three model-specific pricing decisions

Flash is non-thinking mode, so the schedule and accepted-answer boundary stay separate from provider-wide scheduling and the broad Flash/Pro comparison.

1. UTC peak/off-peak plus cache-hit monthly schedule

UTC trafficCache statusFormulaDecision
Peak allocation pHit/miss separatep × peak + (1 − p) × off-peakMeasure p in UTC
Off-peak allocation 1 − pHit/miss separate100K calls × applicable input/output ratesSchedule flexible work
Dated rateModel price recordCurrent Flash rate onlyProvider schedule source required for numeric discount

2. Flash versus Pro cost-per-accepted-answer boundary

ModeToken cost / 100KAccepted-answer rateBoundary
DeepSeek V4 Flash$237.60UnavailableThinking premium cannot be inferred
DeepSeek V4 Pro$712.80UnavailableRecord accepted answers before switching

3. Operational sensitivity

VariableDated valueKeep separate from token price
Retry rateUnavailableMultiply full request cost when measured
LatencyUnavailableDo not convert milliseconds to dollars
SLA / availabilityUnavailableNo SLA claim in pricing record

Verified 2026-08-14. Luna is the data owner for this rendered decision module. “Unavailable” means the current dated registry has no model-specific evidence; it is not a zero. First-party price source · Run this scenario in the playground.

All results are server-rendered for DeepSeek V4 Flash; formulas expose fixed inputs and missing evidence remains visibly unavailable.

Batch 62 · exact-model pricing decision contributions · verified 2026-09-07

Exact model boundary: DeepSeek DeepSeek V4 Flash (deepseek-v4-flash). Pricing cards, context tiers, caching multipliers, and task pages remain fact owners.

Prompt cache hit versus cache miss cost ledger

Frozen Batch 62 scenario board. Formula / deterministic rule: cost = (cache_miss_in * 0.44 + cache_hit_in * 0.11 + out * 1.32) / 1M; cache hit saves 75% Boundary: Owns DeepSeek V4 Flash prompt caching economics.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch62-deepseek-v4-flash-m1-r1
100% cache miss cold prompt
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=100% cache miss cold prompt; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 100% cache miss cold prompt is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m1-r2
50% cache hit warm prompt
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=50% cache hit warm prompt; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 50% cache hit warm prompt is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m1-r3
80% cache hit enterprise system prompt
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=80% cache hit enterprise system prompt; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 80% cache hit enterprise system prompt is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m1-r4
95% cache hit document Q&A loop
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=95% cache hit document Q&A loop; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 95% cache hit document Q&A loop is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m1-r5
cache eviction on cold start
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=cache eviction on cold start; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — cache eviction on cold start is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m1-r6
unsupported multi-part schema
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=unsupported multi-part schema; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unsupported multi-part schema has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

High-volume classification and ETL pipeline budgeting

Frozen Batch 62 scenario board. Formula / deterministic rule: pipeline_cost = records * ((doc_tokens * 0.44 + json_out * 1.32) / 1M) Boundary: Owns high-volume data pipeline economics for DeepSeek V4 Flash.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch62-deepseek-v4-flash-m2-r1
100K structured web scrape records
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=100K structured web scrape records; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 100K structured web scrape records is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m2-r2
500K customer sentiment reviews
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=500K customer sentiment reviews; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 500K customer sentiment reviews is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m2-r3
1M log classification events
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=1M log classification events; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 1M log classification events is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m2-r4
10M enterprise data enrichment batch
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=10M enterprise data enrichment batch; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 10M enterprise data enrichment batch is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m2-r5
rate limit throttling backup
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=rate limit throttling backup; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — rate limit throttling backup is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m2-r6
unparseable JSON output retry
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=unparseable JSON output retry; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unparseable JSON output retry is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: DeepSeek API documentation; verification date 2026-09-07. Missing or conflicting joins fail closed.

DeepSeek V4 Flash vs Pro thinking mode economic trade-off

Frozen Batch 62 scenario board. Formula / deterministic rule: cost_ratio = pro_cost / flash_cost; evaluates when reasoning tokens justify cost Boundary: Owns Flash vs Pro model routing decision boundaries.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch62-deepseek-v4-flash-m3-r1
simple classification (Flash optimal)
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=simple classification (Flash optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — simple classification (Flash optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m3-r2
data formatting & translation (Flash optimal)
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=data formatting & translation (Flash optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — data formatting & translation (Flash optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m3-r3
multi-step logical reasoning (Pro optimal)
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=multi-step logical reasoning (Pro optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — multi-step logical reasoning (Pro optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m3-r4
complex math verification (Pro optimal)
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=complex math verification (Pro optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — complex math verification (Pro optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m3-r5
hybrid router: 80% Flash / 20% Pro
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=hybrid router: 80% Flash / 20% Pro; workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — hybrid router: 80% Flash / 20% Pro is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m3-r6
untested reasoning requirement
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=untested reasoning requirement; workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — untested reasoning requirement is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the DeepSeek V4 Flash Batch 62 scenario →

Continuous SEO Builder · Batch 73 Audit · 2026-09-08Owner: deepseek-v4-flash

DeepSeek V4 Flash API Pricing: Ultra-Low-Cost High-Speed Intelligence

DeepSeek V4 Flash costs $0.44 per million input tokens and $1.32 per million output tokens ($0.66/M blended at 3:1). A high-efficiency open-weights model designed for rapid classification, document summarization, and large-scale data structuring. Verified 2026-09-08.

Module 1 · DeepSeek V4 Flash Micro-Cost Token Economics
Blended Cost = (Input Tokens × $0.44 + Output Tokens × $1.32) / 1,000,000

DeepSeek V4 Flash delivers high-throughput utility inference at $0.66/M blended tokens.

Boundary: Standard pay-as-you-go rate card; off-peak window and prompt caching discounts apply.
ScenarioRendered Evidence & Bounds
Scenario 1Customer ticket intent classification (800 in, 50 out): $0.000418 per ticket
Scenario 2E-commerce product specification extraction (1.5K in, 200 out): $0.000924 per product
Scenario 3Customer feedback sentiment scoring (2K in, 100 out): $0.001012 per review
Scenario 4Document summary paragraph generation (4K in, 300 out): $0.002156 per document
Scenario 5High-volume webhook data normalization (1K in, 100 out): $0.000572 per webhook
Scenario 6Monthly 100M token classification fleet: $66.00 total infrastructure spend
Module 2 · DeepSeek V4 Flash Off-Peak Tariff & Cache Amortization
Discounted Cost = (Cached Input × $0.044 + Off-Peak Tokens × 0.50 + Generation × $1.32) / 1,000,000

Combining off-peak execution with prompt caching delivers industry-leading data enrichment economics.

Boundary: Evaluates 90% prompt caching discount and 50% off-peak tariff reduction during off-peak hours.
ScenarioRendered Evidence & Bounds
Scenario 1Off-peak asynchronous data tagging: cuts baseline token prices by exactly 50%
Scenario 2Shared JSON extraction schema cache (8K tokens): 82% input cost reduction
Scenario 3Combined off-peak + prompt cache: effective blended rate drops below $0.25/M
Scenario 4Zero cache retention storage fees charged during active continuous sessions
Scenario 5Enables massive bulk data enrichment across multi-million record databases
Scenario 6Reduces enterprise data structuring operational costs by over 70%
Module 3 · DeepSeek V4 Flash High-Volume Batch Extraction Pipeline
Batch Efficiency = Throughput (tokens/sec) / Blended Price ($/M)

Exceptional token throughput pairs with rock-bottom pricing for high-volume enterprise ETL.

Boundary: Evaluates throughput optimization and concurrency scaling for automated enterprise workflows.
ScenarioRendered Evidence & Bounds
Scenario 1Processes 10,000 customer survey responses for under $5.00 total API cost
Scenario 2High streaming throughput (120+ tps) ensures zero queue delays on API gateways
Scenario 3Strict JSON mode adherence guarantees zero downstream serialization pipeline errors
Scenario 4Low-memory footprint supports massive concurrent connection limits on shared infrastructure
Scenario 5Reliable instruction following on complex multi-field schema extraction tasks
Scenario 6Ideal operational choice for high-volume ETL pipelines and real-time content moderation
Explore Related Analyses:DeepSeek provider profileCompare vs DeepSeek V4 ProCompare vs GPT-5.4 NanoCheapest AI API comparison
Batch 84 Exact-Model Pricing ContributionsOwner: deepseek-v4-flash (deepseek)verifiedAt=2026-08-14; revalidation required before current claims.

DeepSeek V4 Flash pricing evidence

Source-backed dated rate shape: dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M. DeepSeek V4 Flash rate, off-peak, and ETL arithmetic only; DeepSeek policy, Pro reasoning, and cheapest-model rankings retain their owners. Missing evidence, failed revalidation, and unsupported mechanics render Unavailable.

Module 1 of 3: Extraction concurrency bill ladder

Novel contribution boundary: Owns dated Flash peak-rate arithmetic for fixed extraction shapes; concurrency and throughput are not asserted. Formula / deterministic rule: spend = requests × (inputTokens × inputCostPer1k + outputTokens × outputCostPer1k) / 1,000

Scenario / field IDExact model, provider, dated rates, and fixed inputsResultState
batch84-deepseek-v4-flash-m1-r1
100K compact ETL requests
model=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=requests=100,000; inputTokens=500; outputTokens=150; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
batch84-deepseek-v4-flash-m1-r2
1M structured extraction requests
model=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=requests=1,000,000; inputTokens=1,000; outputTokens=300; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
batch84-deepseek-v4-flash-m1-r3
10M high-volume rows
model=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=requests=10,000,000; inputTokens=2,000; outputTokens=500; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://api-docs.deepseek.com/quick_start/pricing; registry owner=deepseek; verifiedAt=2026-08-14; freshness gate=2026-09-14. Revalidate before any current-price claim.

Module 2 of 3: Flash→Pro thinking-output premium

Novel contribution boundary: Owns dated model-tier arithmetic with explicit retry input; thinking uplift and accepted quality are not sourced. Formula / deterministic rule: requiredAcceptedRate = ProAttemptCost / FlashAttemptCost; observed acceptance = Unavailable

Scenario / field IDExact model, provider, dated rates, and fixed inputsResultState
batch84-deepseek-v4-flash-m2-r1
500-token ETL pass; 10% retry share; 100 retry-input tokens
model=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=requests=1; inputTokens=500; outputTokens=150; retryInputTokens=100; retryShare=10%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
batch84-deepseek-v4-flash-m2-r2
1K-token extraction; 20% retry share; 250 retry-input tokens
model=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=requests=1; inputTokens=1,000; outputTokens=300; retryInputTokens=250; retryShare=20%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
batch84-deepseek-v4-flash-m2-r3
2K-token long row; 30% retry share; 500 retry-input tokens
model=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=requests=1; inputTokens=2,000; outputTokens=500; retryInputTokens=500; retryShare=30%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://api-docs.deepseek.com/quick_start/pricing; registry owner=deepseek; verifiedAt=2026-08-14; freshness gate=2026-09-14. Revalidate before any current-price claim.

Module 3 of 3: UTC and cache evidence ledger

Novel contribution boundary: Owns exact-model DeepSeek registry evidence only; cache TTL, quota, and off-peak eligibility mechanics are not inferred. Formula / deterministic rule: evidence = exact-model registry field when present; cache TTL, quota, and UTC eligibility = Unavailable

Scenario / field IDExact model, provider, dated rates, and fixed inputsResultState
batch84-deepseek-v4-flash-m3-r1
UTC off-peak window
model=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=evidenceField=off-peak; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
batch84-deepseek-v4-flash-m3-r2
Prompt-cache input rate
model=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=evidenceField=cache; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
batch84-deepseek-v4-flash-m3-r3
Cache TTL or quota mechanic
model=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=evidenceField=ttl; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-08-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://api-docs.deepseek.com/quick_start/pricing; registry owner=deepseek; verifiedAt=2026-08-14; freshness gate=2026-09-14. Revalidate before any current-price claim.

Try DeepSeek V4 Flash pricing analysis →

How fast is DeepSeek V4 Flash?

Tokens / sec
132
TTFT
280 ms
Rank
#10 of 31
$ / M ÷ t/s
$0.0050
Measured with 5 runs on a fixed prompt — see the full methodology.

How much does DeepSeek V4 Flash cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.07
1,000,000$0.66
10,000,000$6.60
100,000,000$66.00

How does DeepSeek V4 Flash compare with other models?

DeepSeek V4 Pro$1.98/MGPT-5 Mini$0.69/MMistral Large 3$0.75/MGemini 3.1 Flash Lite$0.56/M
See all DeepSeek models →

What is DeepSeek V4 Flash best for?

#7 for Summarization#9 for Writing & Content#10 for Long Documents & RAG
Looking for a cheaper option?
Muse Spark 1.3 Contributor is 81.1% cheaper — a config migration. See all 8 alternatives to DeepSeek V4 Flash

Which DeepSeek V4 Flash head-to-head comparisons are available?

DeepSeek V4 Flash vs DeepSeek V4 Pro

What are common questions about DeepSeek V4 Flash?

Is DeepSeek V4 Flash cheaper than GPT-5 Mini?

DeepSeek V4 Flash costs $0.66/M blended tokens, GPT-5 Mini costs $0.69/M — DeepSeek V4 Flash is cheaper.

How much does 1 million tokens cost with DeepSeek V4 Flash?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.66. Pure input costs $0.44/M; pure output costs $1.32/M.

What does DeepSeek V4 Flash cost at high volume?

At 100 million blended tokens a month, DeepSeek V4 Flash costs approximately $66.00. See the cost-at-scale table below for other volumes.

Try DeepSeek V4 Flash for free

Run real prompts against DeepSeek V4 Flash and every other model on this page in one workspace.

Try DeepSeek V4 Flash Free