DeepSeek gave deepseek-v4-pro users four days. On September 14, 2026 at 04:00 UTC (12:00 Beijing), every request that names deepseek-v4-pro gets routed to DeepSeek-V4.1-Flash and billed at V4.1-Flash rates. Nothing errors. Nothing warns. Your bill drops, and a different model starts answering your prompts.
That last part is the one to plan for. The release note frames the change as an upgrade, and on DeepSeek’s own numbers it is. But “silently rerouted” is not the same as “tested”. If your product depends on a particular tool-call shape or a latency budget you tuned against V4-Pro, you want to know what changes before Sunday, not after.
This guide covers what DeepSeek announced, what happens if you do nothing, a migration checklist, and a way to build a Pro versus Flash regression suite in Apidog so the cutover is a non-event. For the model itself, read what DeepSeek-V4.1-Flash is first.
TL;DR
- Cutover: September 14, 2026, 04:00 UTC.
deepseek-v4-prorequests go to V4.1-Flash from that moment. - Your calls won’t fail and you won’t pay Pro prices. Peak cache-miss input drops from $1.32 to $0.30 per 1M tokens (77% less); output drops from $3.96 to $1.20 (70% less).
- Rename the model id to
deepseek-flashyourself, then re-test function calling, structured output, streaming, and reasoning effort. - DeepSeek says V4.1-Flash beats V4-Pro on every listed benchmark. Those are vendor numbers. Run your own evals.
What DeepSeek announced on September 10
The September 10 changelog entry covers two things: V4.1-Flash going GA on the API, and V4-Pro retiring four days later.

On V4-Pro, the wording is direct. DeepSeek says V4.1-Flash “has comprehensively surpassed V4 Pro in performance, cost, speed, and total time”, citing “tests by multiple parties”. From September 14 at 04:00 UTC, requests to deepseek-v4-pro are served by V4.1-Flash and priced as V4.1-Flash. There is no grace period where Pro keeps running under a legacy name.
The naming cleanup reaches past Pro:
| Model name | Status after September 10 |
|---|---|
deepseek-flash |
New canonical id for V4.1-Flash |
deepseek-v4-flash |
Still accepted, served by V4.1-Flash |
deepseek-v4-flash-vision-exp |
Still accepted, served by V4.1-Flash |
deepseek-v4-pro |
Served by V4-Pro until September 14 04:00 UTC, then rerouted to V4.1-Flash |
V4-Flash and V4-Flash-Vision-Exp as models are retired; only their names live on as aliases. The base URLs don’t move: https://api.deepseek.com for the OpenAI-compatible format and https://api.deepseek.com/anthropic for the Anthropic-compatible one. The how-to for the V4.1-Flash API covers the new model id, reasoning control, and image input in detail.
What happens if you do nothing
Short version: your integration keeps working and gets cheaper. Longer version: five things change under you.
A different model answers. V4.1-Flash is a 552B-parameter MoE with a new Causal Encoder-Decoder layout: 40 layers, 20 encoder and 20 decoder, 8B parameters active during prefill and 16B during decode. Your prompts now hit a different active-parameter budget and a different attention design (Compressed Sparse Attention 2). Expect different phrasing, different default verbosity, and occasionally different decisions on borderline tool calls.
The output ceiling is the same, the recommended settings aren’t. Both models list 1M context and 384K max output. The V4.1-Flash model card recommends temperature 1.0, top_p 0.95 or 1.0, and max tokens of 256K or more. If you set a tight max_tokens on V4-Pro to cap cost, check whether reasoning output now truncates before the answer arrives.
Reasoning effort works on a different scale. DeepSeek describes V4.1-Flash reasoning effort as “continuously controllable” on a 1 to 100 scale. V4-Flash accepted reasoning_effort plus extra_body={"thinking": {"type": "enabled"}}. How the 1 to 100 scale maps onto the API parameter is [VERIFY] against the current docs; don’t assume a named level like "high" means the depth it did on Pro.
Pricing changes in your favor, with a peak window. Peak hours are Monday to Friday, 01:00 to 04:00 and 06:00 to 10:00 UTC. Off-peak is half of peak. Rerouted Pro traffic pays Flash rates in both windows.
Concurrency headroom jumps. V4-Pro’s limit was 500 concurrent requests. Flash’s is 2,500. Rerouted traffic inherits the higher limit, so revisit any client-side backoff tuned to 500.
Migration checklist
Do these in order. Together they turn a silent reroute into a deliberate release.
- Rename the model id. Search your codebase and config for
deepseek-v4-proand replace it withdeepseek-flash. Do this even though the alias keeps resolving: explicit ids make incidents easier to read later.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DEEPSEEK_API_KEY"],
base_url="https://api.deepseek.com",
)
response = client.chat.completions.create(
model="deepseek-flash", # was "deepseek-v4-pro"
messages=[
{"role": "system", "content": "You triage support tickets. Return a JSON object with priority, team, and summary."},
{"role": "user", "content": "Customer reports checkout returns 502 after applying a discount code."},
],
response_format={"type": "json_object"},
)
print(response.choices[0].message.content)
- Review reasoning effort settings. List every place you set
reasoning_effortor a thinking toggle. Decide per endpoint whether you want depth or latency, then test each against the new scale instead of carrying the old value across. - Re-run function-calling and structured-output tests. Tool-call argument formatting is where model swaps bite hardest. If you wrote tests when you set up function calling on V4-Pro, run them against
deepseek-flashnow. If you didn’t, the Apidog workflow below gives you a set. - Check your streaming parsers. V4.1-Flash streams reasoning deltas and answer deltas separately in thinking mode. Confirm your SSE handler doesn’t concatenate them and that your “first token” timer measures what you think it does.
- Re-baseline latency and cost. Record p50 and p95 latency plus tokens per request on V4-Pro this week, then the same on
deepseek-flash. You’ll want both numbers when someone asks why the dashboard moved on Monday. - Update dashboards and budgets. Cost alerts calibrated to $3.96 output at peak go quiet once you’re paying $1.20, and a budget that never fires is one nobody looks at. Reset thresholds to the rates on the pricing page and fix any per-model breakdown that filters on the old id.
If you call the Anthropic-compatible or Responses API formats instead of chat completions, the same steps apply; the V4-Pro API formats comparison shows what each request looks like so you can map the fields.
V4-Pro vs V4.1-Flash at a glance
Benchmark figures are DeepSeek-reported, from the V4.1-Flash model card. Prices are USD per 1M tokens, effective September 10, 2026.
| deepseek-v4-pro | deepseek-flash (V4.1) | |
|---|---|---|
| Active parameters | Not restated on the V4.1 card | 8B prefill / 16B decode |
| HumanEval | 76.8 | 79.4 |
| GSM8K | 92.6 | 93.0 |
| DeepSWE v1.1 | 62.7 | 74.2 |
| Terminal-Bench 2.1 | 87.9 | 90.6 |
| Input, cache hit (off-peak / peak) | $0.022 / $0.044 | $0.003 / $0.006 |
| Input, cache miss (off-peak / peak) | $0.66 / $1.32 | $0.15 / $0.30 |
| Output (off-peak / peak) | $1.98 / $3.96 | $0.60 / $1.20 |
| Context / max output | 1M / 384K | 1M / 384K |
| Concurrency limit | 500 | 2,500 |
The biggest gap is DeepSWE, up 11.5 points. The smallest is GSM8K, up 0.4. If your workload looks like grade-school math, expect parity; if it looks like multi-file code edits, DeepSeek’s numbers say you gain. The three-way breakdown including V4-Flash is in V4.1-Flash vs V4-Pro vs V4-Flash.
The risk: vendor benchmarks measure vendor tasks
Every number in that table came from DeepSeek. “Tests by multiple parties” is DeepSeek’s phrase, and the parties are not named in the release note. That doesn’t make the numbers wrong. It means they were measured on benchmark suites, not on your prompts, your tool schemas, or your users’ messy inputs.
A model can score higher on Terminal-Bench and still change how it formats a tool-call argument your parser depends on. Cheaper output tokens don’t help if the model writes twice as many. The only way to know is to run your own traffic through both models while both exist. You have until September 14.
Build a regression suite in Apidog before the cutover
Here’s a workflow in Apidog that gives you a repeatable Pro versus Flash comparison and hands the run to CI.
- Import your existing V4-Pro requests. Import an OpenAI-compatible OpenAPI spec, or paste the curl commands you use in production. Put
DEEPSEEK_API_KEYin an environment and reference it asBearer {{DEEPSEEK_API_KEY}}in the Authorization header. Add a second variable,{{MODEL_ID}}, set todeepseek-v4-pro. - Duplicate each request with
deepseek-flash. Set{{MODEL_ID}}todeepseek-flashin a second environment, or hardcode the two ids in paired requests. Same prompts, same tools array, samemax_tokens. The only difference should be the model. - Add assertions on JSON shape and tool-call structure. For structured output, assert that
choices[0].message.contentparses as JSON and contains the keys you expect. For function calling, assert thatchoices[0].message.tool_calls[0].function.nameequals the tool you expect and thatargumentsparses. These checks catch silent format drift. - Run both as one test scenario. Chain the Pro and Flash requests into a single scenario so each run produces one report. Include a streaming request with
stream: true; Apidog renders SSE events one by one, so you can see whether reasoning and answer deltas arrive the way your parser expects. - Diff the results. Compare the two halves of the report: assertion pass rate, response time, and
usage.completion_tokens. A model that passes every assertion but emits 40% more output tokens changes your cost math, and you’ll see it before the bill does. - Schedule it in CI with
apidog-cli. Run the scenario from your pipeline daily through September 14 and keep it running after. Once Pro is gone, the Pro half returns Flash responses too and the diff collapses to zero: confirmation that the cutover happened and your assertions still hold.
Download Apidog to set this up. The flow lives in a shared project, so whoever owns the bill and whoever owns the prompts read the same report.
FAQ
Will my V4-Pro calls fail on September 14? No. Requests that specify deepseek-v4-pro are rerouted to V4.1-Flash. You get a 200 response and a V4.1-Flash answer. The failure mode is behavioral, not an error code.
Do I get charged V4-Pro prices after the cutover? No. Rerouted requests are billed at V4.1-Flash rates: at peak, $0.30 instead of $1.32 per 1M cache-miss input tokens and $1.20 instead of $3.96 per 1M output tokens.
Should I rename the id to deepseek-flash or keep deepseek-v4-pro? Rename. The alias works, but an explicit id means your logs, dashboards, and cost reports say what’s being called.
Does the V4-Pro-0813 setup still apply? The key, base URL, and request shape from the V4-Pro-0813 API guide carry over unchanged. What changes is the model behind them and the reasoning-effort control, so re-test instead of assuming.
Is V4.1-Flash worse than V4-Pro at anything? DeepSeek’s published benchmarks show gains on every listed task. Independent results weren’t available at publish time. Treat “worse at nothing” as a claim to test.
Four days is enough
The migration is one string change plus a test run. The string change takes a minute. The test run tells you whether the model you’ll pay for next week behaves like the one you designed around. Build the suite, run it while deepseek-v4-pro still resolves to Pro, and read the diff. If it’s clean, you get a 70% cheaper bill and five times the concurrency for free. If it isn’t, you found out on your schedule, not DeepSeek’s.



