On August 6, DeepSeek dropped a short announcement that sent ripples through the developer community: they’re planning a “significant” price increase for their API services. No specific numbers yet, just a warning that the hike is coming and it’s going to be big.
Honestly? I wasn’t surprised.
Just three weeks ago, DeepSeek introduced peak-hour pricing—double the cost during weekday business hours. And before that, in May, they made a permanent price cut that cemented their reputation as the industry’s “price butcher”. V4-Pro dropped to about $0.43 per million input tokens and $0.86 for output—a 75% reduction from the original.
So what changed?
The numbers tell the story.
On July 31, DeepSeek released the official version of V4-Flash. The performance boost was massive. According to Artificial Analysis, V4-Flash-0731 scores 50 on their intelligence index while costing about 3 cents per task. For comparison, Anthropic’s Claude Fable 5 averages $3.15 per task, and GPT-5.6 Sol runs about $1.86. That’s not a gap—that’s a chasm.
Then the usage numbers came in. OpenRouter data shows V4-Flash hit 7.22 trillion tokens in weekly calls, taking the global top spot. On August 1 alone, OpenCode reported 8 trillion tokens processed in a single day—5 trillion from free trials and 3 trillion from paid usage. By August 5, that daily number had climbed to 7.5 trillion, with V4-Flash accounting for 6.3 trillion—84% of the total.
Let that sink in. One model, eating 84% of the traffic on one of the biggest coding platforms.
And here’s the kicker: V4-Flash is stupid cheap. The token price is roughly 1/12 of V4-Pro, 1/15 of Zhipu AI’s GLM 5.2, and 1/53 of Kimi K3. At around $0.14 per million input tokens and $0.28 per million output, running a typical agent session costs about $0.13—less than a dollar.
People aren’t even thinking about optimization anymore. Why would they? At that price, you just throw tokens at the problem.

But here’s the problem with that.
Agent workloads don’t care about peak hours. They run batch jobs, async tasks, across time zones. The peak-valley pricing DeepSeek introduced in mid-July was supposed to smooth out demand, but it barely touches agent traffic. And agents are exactly where the token consumption is exploding.
DeepSeek’s infrastructure is getting hammered. On August 4, the API suffered multiple performance degradation events. The official service status page shows fault rates spiking after the Flash release.
The math is brutal. DeepSeek’s pricing philosophy was always “ten months to recover hardware costs”. But when usage grows this fast, that math breaks. You can’t keep selling compute below cost when the compute bill is doubling every week.
What’s the increase actually look like?
My guess? They’re rolling back to pre-discount prices. The May “permanent discount” brought prices to one-quarter of the original. A return to full price would be a 300% increase. “Significant” is probably an understatement.
And it’s not just about Flash. V4-Pro official is coming soon, probably alongside DeepSeek’s new Harness coding agent tool. That’s going to drive even more usage. If they don’t raise prices now, everything crashes.
The bigger picture here.
This isn’t just DeepSeek. The entire industry is hiking prices. Zhipu AI has raised prices multiple times this year. Tencent’s Hunyuan saw increases up to 460%. The era of subsidized AI is ending.
DeepSeek’s pivot from “price butcher” to sustainable business is actually healthy. You can’t burn money forever. The question is whether the increases will kill the developer ecosystem that grew up around these cheap APIs.
For now, I’m not panicking. Even at double the current price, V4-Flash would still be competitive. The performance-per-dollar ratio is just that good.
But I’m also not going to pretend this doesn’t matter. If you’re building on DeepSeek’s API, start optimizing now. Cache aggressively. Schedule non-urgent tasks off-peak. And maybe don’t blow through tokens like they’re free.
Because they’re about to cost a lot more.