The cheapest frontier-class AI in the world just got more expensive. But the test that made DeepSeek’s V4 Flash a star got harder. DeepSeek said it is raising prices for its V4 Flash and V4 Pro models, and they did exactly that.
Depending on the model, the token type and the time of day, the increase runs from a little over half again to as much as eleven times the old rate. Cache hits, when the model reuses a prompt instead of starting fresh, rise between 52% and 1,100% . One thousand and one hundred percent, and still, it’s a fraction of the western models’ prices, as if they want to mog them intentionally.
The new structure keeps seventeen of every twenty-four hours at half price, a deliberate nudge to move batch work into the cheap window and treat “when you run a task” as part of the cost.
Expensive, but can’t deliver?
The timing is pretty awkward. V4 Flash has topped usage leaderboards and drawn “total monster” praise since its July 31 beta, largely because it delivered near-frontier quality at a fraction of the price. Then Composio ran it through eight agent harnesses, Claude Code, Codex, OpenCode and others, on thirty deliberately hard, multi-step tasks spanning Gmail, GitHub, Slack and Google Sheets.
Of 240 runs, 129 passed, and only six of the thirty workflows were finished by every harness. Same model, sharply different results depending on which stack it ran on, and this brings a brand new, rather inconvenient question that must be factored in by anyone who works with AI. Is cheap actually cheap?
Because if you have to do twice as much work when your workflows fail again and again, the cheap becomes not that cheap at all very fast. You also spend more time on that. Sounds like a trade-off we would avoid if possible.
There is a new benchmark in our side
Fair to say, that gap is the bigger story, more than the sticker shock. Also, it moves the question from “which model is best” to “which setup makes a model reliable,” and it hands DeepSeek’s rivals a chance to argue that raw benchmark scores do not predict the bill or the outcome. Because you know, the real benchmark is whether they can get the job done or not. And yeah, there is a grain of truth in this argument.
I am not a coder, the coding benchmarks mean a little to me, and if my agents cannot succeed in the tasks that are important for my work, it is very little relief that they are brilliant in other fields. This is the benchmark of the pudding, the real life test, by the real users.
DeepSeek is still far cheaper than OpenAI, Anthropic, Google, xAI and the rest, and analysts stress it can defend its business case for many workloads. But the pure price story is probably over, and the company is now pricing like an incumbent. Peak and off-peak rates, volume economics through caching, and a clear signal that its home market pays the most. This is how business works.
Developers are not running a charity. When you use one of the most sought-after resources of the world (aka computing power), you’d better prepare that the competition will arrive on your turf too, and someone has to pay.
They develop, we build, but for actual synergy everyone must win
If you build on any of these models, being a side project, a startup, a company tool, or whatever, your costs just became a scheduling problem. “Cheap and smart” is no longer a safe assumption for agent work. It’s easy to say that dev companies are building what they think is demanded, but the demand shows up in usage, subscriptions, and user growth, not the other way around, and measuring whether something brings more revenue or not isn’t a difficult metric.
DeepSeek has the right to raise prices based on costs, demand, and even gut feeling. But in this case, the product had better work as intended and expected.
For the rest of us, the point is straightforward. Pick a vendor on reliability for the specific task, not on the headline price, because the era of buying AI on price alone is ending.










