Skip to main content

folkfox

Skip to main content
Skip to content
AI CONSULTANCY

DeepSeek V4 Just Released a Model Named After Its Own Expiry Date

DeepSeek quietly served a two day beta of deepseek v4, its fastest model family, under an ID that announces its own death: deepseek-v4.1-flash-expires-on-0910. The speed is real, the price cut is real, and the small print is where the story lives.

Quick answerdeepseek v4.1 flash is a two day beta of DeepSeek's fastest model, served through the deepseek api until 10 September 2026, with community measured speeds around 400 tokens per second and a price cut landing on launch day.
SECTION 01

The deepseek v4 beta that carries its own expiry date#

The deepseek v4 story took a strange turn this week, and it is written in the naming. A fox reads the small print before it reads the headline, because the small print is where the trap or the treat lives. On 8 September 2026, DeepSeek began quietly serving an intermediate checkpoint of its fastest model through its existing API, under a model identifier that announces its own funeral: deepseek-v4.1-flash-expires-on-0910. No new endpoint, no new key, just a name change in the request body, and roughly two days to test it before the route goes dark around 10 September.

The beta appeared the way DeepSeek tests everything, and access details spread through community posts reproduced by DeepSeekV4Pro's beta report, then through a notice in its official community groups, picked up by Pandaily and dissected by OrcaRouter's beta report. The notice promises a new architecture with native multimodal support, stronger capability, faster responses and lower cost. What it does not include is a model card, a benchmark table, a context window specification or a multimodal request schema. The claims describe a direction, not a published specification.

The billing is clever: the beta runs at existing deepseek v4 Flash rates while it lasts, and each account is capped at 20 concurrent requests, against 2,500 for the production Flash model. That cap is the tell. A 20 request ceiling is an evaluation lane, not a production highway, and any buyer who missed it would learn the difference at scale.

So the first honest read is this: deepseek v4 is entering its theatre phase, deepseek v4.1 Flash is a marketing-grade teaser wrapped in an engineering-grade artefact. The model is real, the route is real, and the expiry is stamped into the name itself, which is either refreshing honesty or extremely disciplined hype, depending on what launches on 10 September.

SECTION 02

The speed is real, the benchmark whisper is not yet#

Speed is the one claim of this beta that outside testers could verify immediately, and they did. Community measurements from the beta window put decoding speed between 355 and 427 tokens per second, with time to first token around 178 milliseconds, as ByteIota's beta analysis summarises. A 1,500 token code response completes in about 4.4 seconds. Against the current deepseek v4 Pro serving numbers of roughly 63 tokens per second and 766 milliseconds to first token, that is a sixfold throughput jump and a 77 per cent cut in latency.

Decoding speed, measured by testers
Bar chart comparing decoding speed of DeepSeek V4 Pro at 63 tokens per second with V4.1 Flash at 400 tokens per secondV4 Pro: 63V4.1 Flash: 400400 tok/s300 tok/s200 tok/s100 tok/s0 tok/s63 tok/sV4 Pro400 tok/sV4.1 Flash
Bar chart comparing decoding speed of DeepSeek V4 Pro at 63 tokens per second with V4.1 Flash at 400 tokens per second
ItemValue
V4 Pro63
V4.1 Flash400
Community tests clocked V4.1 Flash decoding at roughly 400 tokens per second against about 63 for V4 Pro, a sixfold jump on the same API.

Independent video tests echoed the same shape, one public test recording around 400 tokens per second on deepseek v4.1 Flash, as this hands-on test showed, and production deepseek v4 Flash numbers sit documented in the 7MinAI usage guide, decoding around 300 to 400 tokens per second on real tasks, with one tester generating a 24,000 token landing page in about two minutes. That is the kind of number that makes a fast llm meaningful to anyone running agent loops, where latency compounds across every tool call.

The whisper nobody has sourced is the headline number doing the rounds, that the beta hits 98 per cent of GPT-6 Astra's score on design benchmarks at 1.4 per cent of the cost. No model card, no benchmark table and no DeepSeek publication supports it yet, and until one does it belongs in the rumour column, not the scoreboard. The honest fox follows the scent only as far as the track exists.

SECTION 03

A price cut, a peak hour trap, and a silent reroute#

Money is where this story stops being a curiosity. On 10 September at 12:00 Beijing time, DeepSeek cuts its Flash series pricing, with cache hit prices falling up to 60 per cent, cache miss prices around 33 per cent and output prices around 11 per cent, as Gate News reported. The same adjustment introduces peak hour rates at double the off-peak price, a two speed tariff that changes how you read any llm pricing quote.

DeepSeek API rates, dollars per million tokens
V4 Pro output
$0.87
V4 Pro input
$0.435
V4 Flash output
$0.28
V4 Flash input
$0.14
Production rate card as published: Flash input at $0.14, Flash output at $0.28, Pro input at $0.435 and Pro output at $0.87.

The numbers above are the published card on the DeepSeek API pricing documentation as of July 2026: Flash input at 14 cents, Flash output at 28 cents, Pro input at 43.5 cents and Pro output at 87 cents per million tokens. The 10 September cut reshuffles the Flash half of that card, and the peak window runs 09:00 to 12:00 and 14:00 to 18:00 Beijing time, which is 01:00 to 04:00 and 06:00 to 10:00 in London on British Summer Time.

Then there is the quietest and most consequential line in the whole launch. From 12:00 Beijing time on 10 September, all deepseek v4 Pro API requests will reportedly route automatically to V4.1 Flash at Flash pricing, with no code changes required and no published opt-out, as ByteIota reported citing TechNode. A vendor quietly swapping the model behind a stable model ID is a change management event for anyone with validated production workflows, regression suites or compliance obligations, and it deserves the same scrutiny as a price rise.

SECTION 04

The IPO that explains the pace#

Every pricing decision looks different once you know the seller is preparing for a public listing. On 9 September, Reuters reported that DeepSeek has tapped CITIC Securities to prepare an IPO on Shanghai's STAR Market, aiming to begin the listing process this year, with deal size and timetable undecided. The deepseek ipo narrative is not a rumour any more, it is a Reuters sourced plan.

The capital context explains the urgency, and the valuation arithmetic is worth pausing on, as Enterprise AI reported. DeepSeek raised about $7.4 billion in June at a valuation above $50 billion, with founder Liang Wenfeng committing 20 billion yuan himself while Tencent became the largest external shareholder at 10 billion yuan, per International Business Times. A second round discussed at 480 to 500 billion yuan, roughly $71 to $75 billion, was suspended in July after private remarks attributed to Liang circulated, and the IPO looks like the alternative route to the same money.

A company preparing for a Shanghai listing has incentives to show three things: growing API revenue, shrinking unit costs and technical momentum that justifies a $75 billion conversation. A two day beta that demonstrates a sixfold speed jump, a price cut that makes competitors look expensive, and a model family that quietly absorbs its own Pro tier, delivers all three in a single news cycle. That does not make the engineering fake, it makes the theatre explainable, and a buyer who understands the incentive reads the announcements differently.

The talent story runs alongside. DeepSeek has reportedly expanded to 1,000 employees with 150 open positions, while losing staff to better funded rivals including ByteDance and Xiaomi, per Gate News. Fast, cheap models are also a recruitment poster, and the beta's expiry stamp turns next week's launch into an event rather than a changelog.

SECTION 05

Five checks before you let a fast llm near production#

None of this means DeepSeek is doing anything wrong, but the quarry has moved and the old map is useless. It means the purchase decision has changed shape, from a static comparison of model cards to a live assessment of pricing regimes, expiry stamps and rerouting policies. Five checks separate the buyers who benefit from this beta from the ones who get surprised by it.

The first check is pinning the model ID. Treat deepseek-v4.1-flash-expires-on-0910 as a configuration value, never a constant, and keep deepseek-v4-flash as the documented rollback path, exactly as the OrcaRouter report advises. The second is running your own evals, because no model card exists yet and the 98 per cent Astra whisper is unsourced. The third is respecting the 20 request concurrency cap, which marks this as an evaluation lane.

The fourth is auditing the reroute, because a vendor silently moving your Pro traffic to a Flash model is a change control event your compliance team should sign off. The fifth is modelling the peak window and the cache hit ratio, because both move your real bill more than the headline cut.

The five checks, in order
Pin the model ID

Keep the beta ID configurable and deepseek-v4-flash as the rollback path, never hard-coded.

Run your own evals

Test text quality, tool use and multimodal behaviour on your workloads before trusting any headline.

Respect the cap

Treat the 20 concurrent request ceiling as proof this route is for evaluation, not production.

Audit the reroute

Get the V4 Pro to Flash substitution in writing, with regression and compliance sign-off.

Model the real price

Price the peak window in your timezone and your cache hit ratio, not the headline cut.

What the beta claims, what testers measured, and what remains unsourced.
The claimStatusWhat a buyer should do
Around 400 tokens per second decodingCommunity measuredBenchmark it against your own latency budget
Native multimodal supportAnnounced, schema unpublishedTest image input behaviour before assuming it works
98% of GPT-6 Astra at 1.4% of costUnverified whisperIgnore until a model card or benchmark table exists
V4 Pro requests reroute to Flash on 10 SepReported via TechNode, no official docsDemand written confirmation and an opt-out path
New architecture, not an incremental patchVendor claimTreat as context, judge only the outputs
  • Around 400 tokens per second decodingCommunity measuredBenchmark it against your own latency budget
  • Native multimodal supportAnnounced, schema unpublishedTest image input behaviour before assuming it works
  • 98% of GPT-6 Astra at 1.4% of costUnverified whisperIgnore until a model card or benchmark table exists
  • V4 Pro requests reroute to Flash on 10 SepReported via TechNode, no official docsDemand written confirmation and an opt-out path
  • New architecture, not an incremental patchVendor claimTreat as context, judge only the outputs
The beta, in three numbers

Decoding speed

400 tok/s

Community measured peak, roughly six times V4 Pro.

Concurrency cap

20

Requests per account, against 2,500 for production Flash.

Beta window

2 days

Served 8 September, expires on or around 10 September.

The deeper point for any buyer is that llm pricing has stopped being a menu and started being a market, with expiry dated products, time of day tariffs and silent model substitution arriving in the same release. That is a vendor strategy problem until you have a procurement strategy, which is precisely the work folkfox AI consultancy does for businesses in regulated and awkward categories.

When vendors move faster than documentation, the discipline of separating report from finding is a survival skill, the argument running through our interactive story on rumour and disclosure. And when the same week brings a model with an expiry date and a treaty with a four month deadline, the pattern is clear: the infrastructure is speeding up while the governance is still catching its breath, a theme we covered in our superintelligence treaty analysis.

If your team is about to route real traffic at a model whose name ends in expires-on-0910, review folkfox pricing, see how we think about content built on fast inference, or start the conversation before the grains run out. The hourglass on the ridge has very little sand left.

Questions

Frequently asked questions#

What is DeepSeek V4.1 Flash?

DeepSeek V4.1 Flash is an intermediate checkpoint beta of DeepSeek's fastest model, served through the existing DeepSeek API from 8 September 2026 under the model ID deepseek-v4.1-flash-expires-on-0910, with native multimodal support and a new architecture claimed by the vendor.

When does the DeepSeek V4.1 Flash beta expire?

The model ID itself carries the expiry: deepseek-v4.1-flash-expires-on-0910 is expected to stop serving around 10 September 2026. DeepSeek has not published an exact cutoff time or timezone, so testers should keep the production deepseek-v4-flash ID as a rollback path.

How fast is DeepSeek V4.1 Flash?

Community tests measured decoding between roughly 355 and 427 tokens per second with around 178 milliseconds to first token, against about 63 tokens per second and 766 milliseconds for V4 Pro, roughly a sixfold speed improvement.

Does the beta really beat GPT-6 Astra at 1.4 per cent of the cost?

That figure is circulating on social channels but has no published source, no model card and no benchmark table behind it. Until DeepSeek publishes one, it should be treated as unverified rather than fact.

Will DeepSeek V4 Pro API calls automatically switch to V4.1 Flash?

Reported by ByteIota citing TechNode: from 12:00 Beijing time on 10 September 2026, V4 Pro requests will route to V4.1 Flash at Flash pricing with no code changes and no published opt-out. DeepSeek has not confirmed this in official documentation, so buyers should demand written confirmation.

Is DeepSeek really planning an IPO?

Reuters reported on 9 September 2026 that DeepSeek has engaged CITIC Securities to prepare a listing on Shanghai's STAR Market, aiming to begin the process this year, alongside a funding round discussed at roughly $75 billion. Deal size and timetable remain undecided.

Keep reading

Read more on this topic#

Routing real traffic at a model with an expiry date?

folkfox reviews vendor pricing regimes, rerouting policies and evaluation pipelines so your AI spend survives the next surprise launch.

Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.