English · 6 min read
DeepSeek V4.1 Flash: what changed for your business
DeepSeek V4.1 Flash shipped on 10 September 2026 with a new API id, lower prices, and a plan to route V4 Pro traffic to it from 14 September. Here is what a small business should change, and what it can ignore.
By Dapols ·
The short answer: DeepSeek V4.1 Flash replaced V4 Flash on 10 September 2026. The API id is now deepseek-flash. Input got cheaper and output got slightly cheaper. From 14 September 2026, 04:00 UTC, every call to deepseek-v4-pro is served by V4.1 Flash at Flash rates, until a V4.1 Pro launches. DeepSeek has not given a date for that. If you run automations on DeepSeek, change your model id this week and re-test anything that relied on Pro.
Verified 11 September 2026 against DeepSeek's release note, pricing page and Hugging Face model card. Benchmark scores below are DeepSeek's own unless labelled independent.
What shipped
- Release: 10 September 2026, 04:00 UTC. New prices took effect at the same moment.
- API id:
deepseek-flash. The old idsdeepseek-v4-flashanddeepseek-v4-flash-vision-expare retired and temporarily route to V4.1 Flash. - V4 Pro is being phased out. From 14 September 2026, 04:00 UTC,
deepseek-v4-prorequests route to V4.1 Flash and bill at V4.1 Flash rates "until V4.1-Pro launches." No V4.1 Pro exists yet. - Size and limits: 552B-parameter mixture-of-experts, with 8B parameters active on input and 16B on output. 1M-token context, up to 384K output tokens on DeepSeek's API.
- Vision is built in. Image and text in, text out. No separate vision model to pick.
- Open weights under the MIT licence on Hugging Face, so you can also run it on your own hardware.
- On OpenRouter the id is
deepseek/deepseek-v4.1-flash.
The new price
Per 1 million tokens, from DeepSeek's official pricing page:
| Off-peak | Peak | |
|---|---|---|
| Input, cache hit | $0.003 | $0.006 |
| Input, cache miss | $0.15 | $0.30 |
| Output | $0.60 | $1.20 |
Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Every other hour is off-peak, billed at half the peak rate. Those windows have not changed since the 16 August repricing.
Compared with V4 Flash 0731, cache-miss input fell from $0.22 to $0.15 off-peak (about 32% lower) and output fell from $0.66 to $0.60 (about 9% lower). Cache-hit input fell from $0.007 to $0.003.
If you were on V4 Pro, the drop is bigger: $0.66 input and $1.98 output off-peak become $0.15 and $0.60. That saving comes with a trade-off, covered below.
How good is it
DeepSeek's own table, at maximum reasoning effort:
- Agentic coding and terminal work: V4.1 Flash scores 74.2 on DeepSWE v1.1 (V4 Pro: 62.7) and 90.6 on Terminal-Bench 2.1 (V4 Pro: 87.9).
- Pure knowledge and reasoning: it trails V4 Pro, with 90.9 on GPQA Diamond against 92.4.
- Against frontier closed models it is still well behind on the harder coding suites. On Terminal-Bench 3.0 it scores 30.0 against Claude Opus 5's 43.3.
Independent: Artificial Analysis gives it an Intelligence Index of 40 and flags it as very verbose. It writes many more output tokens than the median model, so your real cost per task depends on output length as much as on the per-token price.
What a small business should do
- Change the model id to
deepseek-flash. The old ids route for now, but "temporarily" means they can stop. Update config, not just hope. - If you use
deepseek-v4-pro, re-test before 14 September. Your requests will quietly move to V4.1 Flash. For coding and multi-step automations, you will likely see equal or better results for less money. For knowledge-heavy work, like long research answers or detailed factual Q&A, check a sample of outputs. There is currently no cheap DeepSeek step-up model. - Schedule big batch jobs off-peak. Overnight and weekend runs cost half.
- Watch output length. A verbose model can eat the per-token saving. Cap output tokens on routine jobs like tagging and summaries.
- Keep your team's chat app. DeepSeek saves money on automation volume, not on the ChatGPT or Claude seat people already use every day.
What you can ignore
- Headlines about "V4.1 Pro." It has not launched. Do not plan around it.
- Aggregator prices. Third-party gateways list different rates. Budget against DeepSeek's own page.
Where this fits in a stack
For classification, lead tagging, review summaries and routine drafts, V4.1 Flash is now the cheapest credible option we track. See the DeepSeek tool page for the full price table and DeepSeek vs ChatGPT for the trade-offs. Our earlier note on V4 Pro-0813 is now history.
A Business AI Plan ($99 one-time, optional $29/month to keep it current) names which jobs go to which model for your business. The free 2-minute AI plan finder gives you a first pass.
Sources: DeepSeek release note, 10 Sep 2026, DeepSeek pricing, DeepSeek-V4.1-Flash model card, Artificial Analysis