On July 31, DeepSeek V4-Flash officially launched its API for public testing.
For every million input tokens, it costs $1; for every million output tokens, it costs $2. After a cache hit, the cost for every million input tokens is reduced to just 2 cents. It supports 1 million contexts, thinking modes, tool invocations, and Responses API, with a maximum of 2500 concurrent connections per account.
Next, DeepSeek will implement surge pricing, doubling the price during peak hours. Nevertheless, the price remains surprisingly affordable.
According to a rough conversion provided by DeepSeek, approximately 0.6 tokens are consumed by a single Chinese character. With one million tokens, approximately 1.67 million Chinese characters can be processed. Looking solely at the input cost, processing around 1.67 million Chinese characters would cost only $1.
Over the past year, the large-model industry has become accustomed to longer contexts, more complex reasoning, and increasingly capable Agents. With each model update, the discussion inevitably revolves around who has surpassed whom, how much the benchmark has shifted, and how far we are from AGI.
For AI application entrepreneurs, however, there is a more practical question:
How much will a task actually cost?
The model's capabilities determine what the product can theoretically do, while the pricing decides how much of that capability the team dares to truly incorporate into the product.
This afternoon, I had a lengthy discussion with ChatGPT via voice. I said that the current business model for AI applications essentially boils down to reselling tokens.
It tried several times to come up with a more dignified term for me, such as "retail reasoning capability" or "smart service packaging." I told it not to bother finding euphemisms for this business, as reselling is reselling, and being blunt doesn't make it wrong.
AI applications may seem like the newest entrepreneurial trend in recent years, but fundamentally, they are engaged in a very old business:
Procuring, processing, pricing, and then selling the goods.
And many industrial changes have often begun with a sudden drop in procurement costs.
AI Applications Are Essentially a Business of Reselling Tokens
In the past, people were used to understanding AI applications within the framework of internet products or SaaS.
After software development is complete, it can be distributed to a thousand people or a million people. Adding a new user incurs bandwidth, storage, and service costs, but these costs are usually gradually diluted as the scale increases.
AI applications are different. Every time a user generates an image, writes a report, completes a section of code, or has the Agent perform an additional step, the application company needs to re-invoke the model. Every time a user clicks a button, the upstream provider gets paid again.
Therefore, the basic business model for most pure software AI applications currently in operation is actually quite simple:
Procure tokens from a model company, package the general model capability into writing, programming, design, customer service, video, or Agent products, and then sell them to users through membership, points, and pay-per-use fees.
Take Poe as an example. When launching its API, it directly stated that the pricing target with attached points is to charge based on the price the underlying model supplier charges it.
The upstream company gets paid for the model and tokens, Poe converts them into points on the platform, and then sells them to users along with a unified entry point, payment system, model switching, and development interfaces.

Some companies may say they do not sell tokens but sell results. The customer service software company Intercom designed a price for AI customer service Fin, charging $0.99 for each completed result. User confirmation that the issue has been resolved, no further request for human assistance, or Fin successfully completing a workflow can all be considered as a result.
This is just changing the price tag, charging per seat, per task, or per successfully resolved issue. What changes is the unit of measurement users see. Behind the application company, the model still needs to be invoked, and they still need to calculate how many tokens are required to complete a delivery.
Southern Song Dynasty painter Li Song once painted a picture of "Market Porter Playing with an Infant." A peddler carried a load of goods wandering through the villages and alleys, with pots, bowls, needles, snacks, and toys all hanging on the load. A load of goods is like a mobile small department store, and the peddler earns money from selecting, distributing, and delivering goods.
AI application companies do something similar. Model manufacturers produce intelligence, and application companies integrate this intelligence into specific scenarios. What is bought at the base level is tokens, and what should ultimately be sold is a report, a design, or a video.
Model companies sell intelligence, and application companies sell outcomes.
The profit of this business mainly depends on three price differences.
The first one is the purchase price difference.
For the same model, the actual cost obtained by different teams may not be the same. Companies with a large volume of requests can negotiate discounts, while teams with strong engineering capabilities can leverage caching, asynchronous processing, peak/off-peak usage, and model routing to further reduce the average cost. What others can do for one dollar, you can do for fifty cents, and the remaining fifty cents is profit.
In simpler terms, when purchasing the same item, some people get it at wholesale prices, while others get it at retail prices.
The second one is the usage price difference.
Some people treat AI like a workhorse after purchasing a membership, while others show great ambition when topping up but stop using the product after just two weeks. As long as the average consumption of the entire user base is lower than the estimated usage when designing the package, the unused portion will become a profit margin.
This logic is similar to that of gym memberships and mobile data plans.
However, there is one thing to note here. Profit increases only if upstream procurement is flexible enough and users are underutilizing the service. If the company is locked into a fixed annual contract and the remaining tokens are just sitting in the account, that is not profit but unsold inventory.
Unused portions by users are known as "sedimentation," while unsold inventory by the company is called "overstock."
The third one is the product price difference.
Consuming ten thousand tokens, some products only output a piece of text that requires extensive user modifications, while others have already completed data organization, content generation, fact-checking, format adjustments, and final export. The underlying costs may be similar, but the prices users are willing to pay can be vastly different.
The true value of an AI application depends on this. Procurement determines if this business can survive, while the product determines how much this business is worth.
The resold token is just the base, and the product is the reason for the price difference.
Product Manager, the New Accountant
The trouble is, as AI applications constantly need to be restocked, product issues easily turn into cost issues.
In Manus's official usage instructions, you can already see how this cost shapes the product. Manus deducts points from the monthly subscription based on the task's complexity and the resources needed. When interacting with AI, simple questions are recommended to be handled in Chat mode to avoid initiating a full-blown autonomous Agent unnecessarily; for complex tasks, it is best to check intermediate results first to prevent the Agent from continuing in the wrong direction, thus needlessly consuming more points.

This advice makes perfect sense, of course. However, it also reveals the most challenging aspect of Agent products: the more proactive the model, the harder it is for the product company to predict costs.
A chatbot that stops after one response is relatively easy to account for. An Agent will autonomously break down tasks, search repeatedly, open web pages, call on tools, and then decide on the next steps based on intermediate results. While the user has only given one command, behind the scenes, dozens of model calls may have already taken place.
As a result, products have started teaching users how to minimize the work done by the Agent. The fear is that after crunching numbers for a while, the accountant will start taking over the product. When deciding whether or not to launch a feature, the team's primary consideration is no longer whether the user needs it but rather how much each individual task will cost. Although the model has enhanced capabilities, the application is hesitant to entrust these capabilities entirely to the user.
Of course, the company needs to control costs. No startup can afford to treat the most expensive model as a commodity for users to casually exploit.
Some products have placed a more effective model in a high-priced package, offering regular users a watered-down version that has gone through multiple tiers of throttling. The model's capabilities have not vanished; they have simply been distributed to different users after settling accounts.
Over time, the industry has witnessed a rather awkward inversion.
Users purchase AI applications hoping the model will understand more, think further, and complete more tasks. However, application companies, in order to control costs, must navigate a delicate balance of having the model read less, think less, and do fewer tasks.
While the product promises intelligence, the team's most practiced action has become curtailing that very intelligence.
Furthermore, if prices increase upstream, application companies must cut back; when competitors acquire lower costs, membership fees in the market immediately drop. With the emergence of new affordable models, existing routing and packages need to be recalculated. Although application companies appear to have control over the product, the most critical variable is always held by the upstream.
Over the past two years, what AI application entrepreneurs have lacked is not necessarily ideas but rather intelligence that is robust and affordable enough to be confidently integrated into their products.
DeepSeek Cuts Down on "Intelligence Toll"
The value of DeepSeek lies first and foremost in the fact that it is not merely a model with reduced capabilities created solely for cost reduction. It boasts around 1 million contexts, supports reasoning patterns, and tool invocation.
When DeepSeek's preview version was released in April, it claimed its inferencing capabilities were close to V4-Pro and performed similarly on simple Agent tasks. The official version updated on July 31 did not alter the model's structure and scale but primarily underwent retraining, with several Agent benchmarks now surpassing the preview version of V4-Pro.

This is important.
Cheap models have existed before. The issue is that while it's great to get a good deal on a model, the goods shouldn't always be flawed. What AI applications truly need is intelligence that is as stable as possible below a certain unit price.
At least from the benchmarks of Agent and code released by DeepSeek, V4-Flash has already crossed the "sufficient" line for a batch of high-frequency tasks, with a price low enough for the team to confidently use it for high-frequency calls.
For most ordinary users, large models have likely already reached a turning point.
Ordinary people don't care about benchmarks. They rarely care whether a model is three points or five points ahead in a certain test; they only care about whether the model in their hands can get the job done. When a model can already read materials, reason, operate tools, and deliver results, further improving its capabilities is certainly valuable. But can this additional value still justify a several-fold price increase? That needs to be recalculated.
In the past, model companies could justify higher prices by claiming their models were more powerful. However, this reasoning doesn't hold up well anymore. Where is the value in the higher price? What problems do the additional capabilities solve? How many users actually need them? All these questions need to be addressed candidly.
DeepSeek didn't make all models cheaper; it made high-priced models into something that needs to be explained to users.
What Pinduoduo truly broke through in its early days wasn't just the prices of goods. It showed consumers for the first time that many everyday items could be sold so cheaply. The brand premium, channel costs, and middlemen that used to be taken for granted suddenly needed to find reasons for their existence.
What Pinduoduo broke was the price perception of goods.
What DeepSeek has broken is the price perception of intelligence.
Entrepreneurs Can Finally Focus on Product Development
For the AI application industry, the most important thing next is that entrepreneurs can finally shift their focus back from upstream.
In the past, with rapid updates to large models, application companies were easily led by them. Today integrating longer context, tomorrow switching to a new reasoning model, and the day after showcasing Agent capabilities on the homepage. The product seems to be constantly upgrading, but the roadmap is not their own. What the model manufacturers release is what the application companies demonstrate.
A typical scenario of upstream prosperity and downstream distraction.
Now, foundational models have the opportunity to become a relatively stable production material. Teams no longer need to revolve around model rankings every day, nor consider "adopting the latest model" as the most important product development.
Once the model stabilizes, the product finally runs out of excuses.
Being able to accomplish a task is just the beginning for a model. What the product needs to do is integrate the task into a real process, handling permissions, data, collaboration, review, and failures, so that users don't have to reteach the AI how to work every time. These aspects are not as glamorous as model parameters and are difficult to benchmark. However, whether users are willing to continue paying in the second month often depends on these details.
The value of an AI application lies not in presenting the model to the user but in making the model disappear into the workflow.
When users no longer need to choose a model, dig into prompt suggestions, or worry about how many rounds of inference are behind the task, the application has truly done its job.
At this point, what application companies compete on is no longer who can access the model first, but who understands a task better and who can turn industry-specific unspoken rules into part of the product.
Models can become more powerful overnight, but industry understanding cannot. APIs can be accessed by everyone, but a set of work methods that users truly accept is hard to replicate.
This also gives many seemingly modest businesses a new chance to succeed.
AI startups don't necessarily have to start from the "next-generation entry point," nor does every company have to create a super agent that serves all humanity. It can simply solve one problem in foreign trade inquiries, provide a set of processes for chain stores, or eliminate a segment of repetitive work for a certain group of professionals.
The market doesn't have to be big enough to accommodate everyone. As long as the problem is specific enough, the delivery is stable enough, and the value users receive exceeds the price they pay, this business can survive.
In the past, when models were too expensive, a small market struggled to support a complete product. Teams had to aim for a larger user base, higher average revenue per user, and faster financing to cover the constantly accruing inference costs. Many startups began telling an overly ambitious story before they had found real users.
Now, the story can be smaller.
A small team can initially serve a small group of people, perfect one thing, and then gradually expand.
This is very similar to how many "century-old establishments" used to operate. When they opened, they didn't set out to dominate a whole street; they just focused on welcoming the few tables of customers at the entrance. By cooking good food, offering reasonable prices, and having customers return, their business naturally grew.
The big model industry likes to talk about scale, parameters, and endgames. However, in the end, application startups still need to return to this simple logic.
DeepSeek will not answer the three questions for entrepreneurs: who you serve, what you solve for them, and why they can't switch to another provider.
It simply allows teams to stop spending most of their time negotiating with upstream partners and enables smaller teams that do not have the capability to train large models or qualify for special pricing to obtain a relatively fair entry ticket.
In the past, the model itself was a barrier. Next, models will be accessible to everyone, but the difference lies in how they are used. The truly scarce resources will gradually shift towards users, data, processes, and trust. The closer you are to actual work, the more opportunities you will have.
After lowering the threshold for AI intelligence, DeepSeek has finally transformed AI applications from a model competition back into a product business.

In his "Commentary on the Expensive and Cheap Grains," Chao Cuo of the Western Han Dynasty once wrote: "Grains are expensive while gold and jade are cheap."
In recent years, large models have been more like gold and jade. The parameters have become larger, the rankings higher, displayed under the spotlight at conferences, visible to all as precious, yet few can truly integrate them into their daily work.
What the application layer needs to do is turn gold and jade into grains.
DeepSeek did not prepare the meal for entrepreneurs; it simply lowered the price of rice.
As for who can open the restaurant, it all depends on individual craftsmanship.
-END-
Welcome to join the official BlockBeats community:
Telegram Subscription Group: https://t.me/theblockbeats
Telegram Discussion Group: https://t.me/BlockBeats_App
Official Twitter Account: https://twitter.com/BlockBeatsAsia
