DeepSeek-V4.1-Flash Pricing Explained: Peak vs Off-Peak Rates and Cache-Hit Math

The full USD and CNY price table for deepseek-flash, the percentage cuts versus V4-Flash and the V4-Pro reroute, peak windows converted to your time zone, and two cache-hit examples with the arithmetic shown.

Medy Evrard

10 September 2026

DeepSeek-V4.1-Flash Pricing Explained: Peak vs Off-Peak Rates and Cache-Hit Math

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

DeepSeek-V4.1-Flash went GA on the API on September 10, 2026, and the release note came with a price cut attached. Cache-hit input now costs $0.003 per million tokens off-peak, cache-miss input $0.15, output $0.60. Four days later, on September 14, DeepSeek starts routing every deepseek-v4-pro request to V4.1-Flash and billing it at Flash rates. If you run anything on the DeepSeek API, your bill changes this week whether you touch your code or not.

This post is the pricing half of the story: the full USD and CNY tables, the percentage changes against V4-Flash and V4-Pro, where the peak windows land in your time zone, and two cache-hit examples with the arithmetic shown. For the model itself, start with what DeepSeek-V4.1-Flash is.

One caveat up front. Every benchmark number in this cluster is a DeepSeek-reported figure from the model card. Prices come from the pricing page, verified on September 10. The last section shows how to measure your own spend with Apidog by reading the usage block on every response instead of trusting an estimate.

button

TL;DR

The full price table

All figures are per 1M tokens, effective September 10, 2026 at 04:00 UTC.

deepseek-flash off-peak deepseek-flash peak deepseek-v4-pro off-peak deepseek-v4-pro peak
Input, cache hit $0.003 $0.006 $0.022 $0.044
Input, cache miss $0.15 $0.30 $0.66 $1.32
Output $0.60 $1.20 $1.98 $3.96
Concurrency 2,500 2,500 500 500

The deepseek-v4-pro columns matter for four more days. From September 14 at 04:00 UTC (12:00 Beijing), requests to that model name are served by V4.1-Flash and billed from the first two columns. The V4-Pro retirement and migration guide covers what to test before then.

If your account settles in yuan, the zh-cn pricing page lists deepseek-flash as:

Off-peak (CNY) Peak (CNY)
Input, cache hit ¥0.02 ¥0.04
Input, cache miss ¥1.00 ¥2.00
Output ¥4.00 ¥8.00

Each CNY row is the USD row times about 6.67. Two details the table leaves out: context caching is automatic, with no flag to set, and the model id is now deepseek-flash. The old names deepseek-v4-flash and deepseek-v4-flash-vision-exp still resolve but land on V4.1-Flash at these rates.

What changed versus V4-Flash and V4-Pro

On August 21, 2026, deepseek-v4-flash billed $0.007 / $0.22 / $0.66 off-peak (cache hit / cache miss / output) and $0.014 / $0.44 / $1.32 at peak. The changelog says pricing was “reduced accordingly” with this release. Here is what “accordingly” means, computed from the two lists:

Line item V4-Flash (Aug 21) V4.1-Flash (Sep 10) Cut
Cache hit, off-peak $0.007 $0.003 57%
Cache miss, off-peak $0.22 $0.15 32%
Output, off-peak $0.66 $0.60 9%

The peak rows carry identical percentages, since both lists double at peak. The pattern is deliberate: the steepest cut sits on cache hits, the smallest on output. DeepSeek is pricing for agent loops that re-read a long prefix on every turn, and the FP4 KV cache in V4.1 (890 bytes per token, about a quarter of V4-Flash’s footprint) is what makes a $0.003 hit price sustainable on their side.

For V4-Pro users the numbers are larger, because the reroute happens at Flash rates:

DeepSeek says V4.1-Flash “has comprehensively surpassed V4 Pro in performance, cost, speed, and total time”, a claim to check against your own prompts before the 14th. For the longer arc of the family’s pricing, including the stretch when DeepSeek raised API prices, the DeepSeek V4 API pricing breakdown has the history.

Peak windows in your time zone

Peak applies Monday to Friday only, in two blocks: 01:00 to 04:00 UTC and 06:00 to 10:00 UTC. That is 9:00 to 12:00 and 14:00 to 18:00 Beijing time. Seven peak hours per weekday, 35 per week, so peak covers about 21% of the 168-hour week. Everything else, including all of Saturday and Sunday, bills at the off-peak rate.

Window UTC Beijing (UTC+8) US Pacific (PDT, UTC-7) Central Europe (CEST, UTC+2)
Peak 1 01:00 to 04:00 09:00 to 12:00 18:00 to 21:00 (previous evening) 03:00 to 06:00
Peak 2 06:00 to 10:00 14:00 to 18:00 23:00 to 03:00 (overnight) 08:00 to 12:00

The Pacific and Central Europe columns use summer time. After the clocks change (October 25 in the EU, November 1 in the US), shift both columns one hour earlier. The UTC boundaries never move.

A US West Coast working day never touches a peak block. A Central Europe morning sits inside Peak 2, and a European nightly batch at 03:00 local pays double; moving it to 23:00 CEST (21:00 UTC) puts it back at half price.

Cache-hit math, worked twice

A hit costs 1/50 of a miss ($0.003 versus $0.15), so the ratio of repeated prefix to fresh tokens decides your input cost more than the raw token count does. If prefix caching is new to you, the prompt caching primer explains what counts as a cacheable prefix: identical bytes at the start of the request.

Example 1: an agent loop with a 20K-token system prompt, 1,000 calls at peak.

A support-triage agent carries a 20,000-token system prompt (policies, tool schemas, few-shot examples), reads 2,000 fresh tokens of ticket text per call, and writes 1,000 output tokens. Over 1,000 calls:

  1. Total input is 1,000 × 22,000 = 22M tokens. Total output is 1M tokens.
  2. The first call misses on all 22,000 tokens. Every later call hits on the 20,000-token prefix and misses on its 2,000 fresh tokens.
  3. Hit tokens: 999 × 20,000 = 19.98M. At $0.006 per 1M: 19.98 × 0.006 = $0.12.
  4. Miss tokens: 22,000 + 999 × 2,000 = 2.02M. At $0.30 per 1M: 2.02 × 0.30 = $0.61.
  5. Output: 1M × $1.20 = $1.20.
  6. Total: $0.12 + $0.61 + $1.20 = $1.93.

With no caching, input alone would be 22 × $0.30 = $6.60 and the job would cost $7.80. Caching removes 75% of the bill. Run the identical loop off-peak and every line halves: $0.06 + $0.30 + $0.60 = $0.96.

The same 1,000 calls on V4-Pro at peak (hits $0.88, misses $2.67, output $3.96) total $7.51, the number a Pro customer should hold next to $1.93 when the reroute lands.

Example 2: a batch job with 10M output tokens, moved off-peak.

A catalog team generates 10,000 product descriptions. Each call sends 1,500 tokens of product data (unique per item, so no cache benefit) and gets back 1,000 tokens of copy. That is 15M input tokens, all misses, and 10M output tokens.

  1. At peak: input 15 × $0.30 = $4.50. Output 10 × $1.20 = $12.00. Total $16.50.
  2. Off-peak: input 15 × $0.15 = $2.25. Output 10 × $0.60 = $6.00. Total $8.25.
  3. Savings from the timestamp alone: $8.25, with no change to the prompts.

Output dominates this job and got the smallest cut in the release, so the off-peak window is the bigger lever. It is a generous window: 15 hours every weekday from 10:00 UTC, plus 48 uninterrupted hours each weekend.

Context, output, and why concurrency is part of the price

Both model names carry a 1M-token context and a 384K-token max output. The model card recommends max tokens of 256K or more, so set that ceiling on purpose: a response that runs to 384K output tokens costs $0.46 at peak, fine for an agent transcript and a bug for autocomplete.

Concurrency is 2,500 for Flash against 500 for Pro. It belongs in a pricing article because the off-peak discount is only worth something if you can push the work through the window, and because deepseek-v4-pro traffic rerouted on September 14 inherits the Flash limit: a Pro integration that used to queue behind 500 slots gets five times the headroom without a config change.

Measure real spend with Apidog

Price tables tell you the rate. The usage object on each response tells you what you were charged for. Here is an Apidog workflow that makes that a repeatable check.

  1. Add the endpoint. Create POST https://api.deepseek.com/chat/completions, or import an OpenAI-compatible OpenAPI spec.
  2. Put prices in environments. Create two Apidog environments, deepseek-peak and deepseek-offpeak, each holding DEEPSEEK_API_KEY plus PRICE_HIT, PRICE_MISS, and PRICE_OUT as {{variables}}. Peak holds 0.006 / 0.30 / 1.20; off-peak holds 0.003 / 0.15 / 0.60.
  3. Send once and read usage. DeepSeek splits input into prompt_cache_hit_tokens and prompt_cache_miss_tokens alongside completion_tokens [VERIFY]. Save the request as a test case.
  4. Build a test scenario that asserts on cache behavior. Chain the saved request twice with the same 20K-token system prompt. On the second step, assert that the hit-token field is at least 19,000 and that computed cost (hits × {{PRICE_HIT}} + misses × {{PRICE_MISS}} + output × {{PRICE_OUT}}, divided by 1M) stays under budget. A miss on step two means the prefix changed, and the assertion catches it before the invoice does.
  5. Compare peak against off-peak. Run the scenario under each environment and diff the cost line, or run it from CI with apidog-cli on a schedule that straddles a UTC peak boundary.

Download Apidog and the whole scenario, environments included, works on the free plan.

Where this leaves your bill

V4.1-Flash is cheaper than V4-Flash on every line and 70 to 86% cheaper than the V4-Pro rates it replaces. Two levers matter: keep your long prefix byte-identical so it hits the cache, and schedule batch work outside the seven weekday peak hours. Then log usage on every call, assert on the hit count, and let the numbers show whether the cut reached your invoice.

Explore more

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, and no shared benchmark

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, and no shared benchmark

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, no shared benchmark. Specs, vendor claims, and how to test both yourself.

30 September 2026

Is GPT-6.1 Sol Free?

Is GPT-6.1 Sol Free?

Is GPT-6.1 Sol free? No: it's for Plus and up in ChatGPT Work and Codex, and the API has no free tier. Here are the cheapest routes, with real cost math.

30 September 2026

How to Use Claude Sonnet 5.5 for Free: Every Route That Works (and the Ones That Don't)

How to Use Claude Sonnet 5.5 for Free: Every Route That Works (and the Ones That Don't)

Is Claude Sonnet 5.5 free? Yes on Claude.ai (web, iOS, Android). Every free route checked, plus what isn't: Claude Code, the API, and Copilot Free.

29 September 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

DeepSeek-V4.1-Flash Pricing Explained: Peak vs Off-Peak Rates and Cache-Hit Math