Editor's Note: Over the past two years, competition in the AI industry has almost always revolved around one core variable: who has the stronger foundation model. GPT, Claude, Gemini, and the subsequently fast-rising DeepSeek, Qwen, and other models have continuously refreshed benchmarks, making "frontier capability" the most important pricing basis for model companies. But as model iteration accelerates and inference costs continue to fall, the discussion is shifting from "whose model is strongest" to "when there are more and more sufficiently good models, how much can the strongest model still be worth." As model capabilities gradually turn from a scarce good into a more easily obtainable foundational capability, a more critical question begins to emerge: what is truly scarce in the AI industry over the long term—the models themselves, or the products, workflows, and distribution capabilities built on top of them?
In the latest episode of the All-In Podcast, tech entrepreneur Jason Calacanis, Social Capital founder Chamath Palihapitiya, Ohalo CEO David Friedberg, and Craft Ventures co-founder David Sacks discussed the recent wave of model releases, the expansion of open-weight models, falling token prices, and the business models of frontier model companies such as Anthropic and OpenAI. Compared with "how much performance the next-generation model has improved," what is more noteworthy about this conversation is that it attempts to answer where the AI industry's profit pool may migrate after model capabilities spread rapidly.

In this conversation, the hosts are in effect breaking down "models getting stronger and stronger" into a set of more fundamental structural questions: when the performance gap between models narrows, why should enterprises still pay a high premium for frontier models? When open-weight models can handle more and more daily workloads, are tokens themselves becoming commoditized? And when the "brain" becomes cheaper and cheaper, will value reconcentrate around agents, enterprise workflows, and vertical applications?
First, model competition is shifting from capability scarcity to cost efficiency. In the past, there was often a clear capability gap between the most advanced models and the second tier, and companies were willing to pay a significant premium to obtain higher accuracy, stronger reasoning, and more stable output. What has changed now is that model releases are becoming more frequent, the performance of open-weight models continues to approach the closed-source frontier, and API providers themselves are constantly cutting prices. This means that for a large number of tasks that do not require extreme capability, "the best" is turning from a necessity into a choice that can be calculated in terms of return on investment. Model performance is still important, but the ratio between price and performance is beginning to matter more than simple benchmark rankings.
Second, tokens may be turning from a high-margin product into a standardized basic input. The logic of the previous round of AI commercialization was relatively straightforward: model companies provide intelligence, developers buy tokens, and then package them externally into applications. But when the same type of task can be migrated across multiple models, and the price gap reaches several times or even dozens of times, companies will naturally assign different tasks to models at different cost tiers. Simple summaries, customer service, internal processes, and ordinary coding tasks can be handed to cheap models, while only a small number of tasks such as scientific research, complex engineering, and high-value financial decision-making continue to call on the strongest models. What is truly being challenged here is not whether frontier models have value, but how much revenue must depend on frontier capability in order to hold up.
Third, the competitive boundary of model companies is being forced to move upward. Chamath likens foundation models to a "brain," and calls the tool-calling, browser, memory, permissions, and task-planning systems built around the model a harness. This distinction is important: when the gap between "brains" keeps narrowing, what ultimately determines the product experience may increasingly be not the model itself, but whether it has "hands and feet" and can actually complete tasks. Agent products like Meta Muse therefore take on another layer of meaning—users do not care which model is called behind the scenes, but whether it can organize email, book flights, process documents, and get work done. The cheaper the foundation model, the easier it is for the economics of such products to hold up.
Fourth, the business model of frontier models may therefore become polarized. At one end is general intelligence that is huge in scale but highly price-sensitive, undertaken by open-weight models and low-cost models; at the other end is Frontier Intelligence, which is smaller in number but willing to pay extremely high premiums, such as complex scientific research, mathematics, engineering, biomedicine, and highly competitive professional scenarios. The key question for companies such as OpenAI and Anthropic is therefore no longer just "can they continue to improve models," but whether they can prove that the incremental value created by the most advanced models is enough to cover their higher training, inference, and capital expenditures over the long term.
If this conversation is compressed into a single judgment, it is this: the core contradiction of AI is gradually shifting from "is intelligence strong enough" to "how is intelligence priced, distributed, and productized."
In this sense, the subject discussed in this article is no longer just OpenAI, Anthropic, or a particular new model, but a repricing of the entire AI industry value chain: as models themselves become increasingly easy to obtain, the truly scarce assets may once again become users, scenarios, workflows, and the ability to convert intelligence into actual value.
The original content is as follows (the original content has been edited for easier reading and comprehension):
TL;DR
Competition among AI models is shifting from "who is strongest" to "who can reach a sufficient level more cheaply," which essentially means model capabilities are beginning to move from scarce goods to standardized supply.
The continued decline in token prices means that the profit margin of foundation models themselves may be compressed, and the AI value chain is shifting from "selling intelligence" to "selling results."
The biggest impact of open-weight models is not replacing frontier models, but taking over a large number of general-purpose workloads, forcing companies to recalculate whether the "premium for the strongest model" is worth it.
In the future, frontier models are more likely to concentrate on high-value complex tasks such as scientific research, engineering, and finance. The key is not covering all scenarios, but proving their irreplaceability.
The more model capabilities converge, the more competitive advantage shifts toward agents, tool calling, memory, permissions, and workflows. Essentially, value shifts from the "brain" to the "execution system."
Enterprise AI cost structures will shift from "uniformly calling the strongest model" to tiered routing, matching different tasks with models at different price and capability levels.
The core business risk for OpenAI and Anthropic is not that models stop improving, but that a large amount of existing revenue is ultimately commoditized by cheaper models.
In the next stage of the AI industry, what is truly scarce may no longer be the models themselves, but users, scenarios, distribution, and the ability to convert intelligence into actual output.
Video Key Points
Models are still getting stronger, but more importantly: they are getting cheaper
Over the past few years, competition among AI models has followed a very clear logic: whoever has the strongest model has scarcity.
Each time GPT, Claude, and Gemini cross a capability threshold, it is enough to bring a new round of product and capital narratives. When the performance gap between models is large enough, enterprises are willing to pay a clear premium for "the strongest intelligence."
But All-In believes this logic is becoming complicated.
On the show, Friedberg listed a series of recent model updates. Although some model names and release dates in the automatic transcription contain errors, the underlying trend is not wrong: model updates are becoming extremely密集, and performance improvements and cost reductions are happening at the same time.
For example, DeepSeek released V4.1 Flash on September 10. This is a 552-billion-parameter MoE (Mixture of Experts) model, but it activates only about 8 billion parameters on input and about 16 billion parameters on output. DeepSeek also lowered its API prices and said the new KV Cache design will significantly reduce VRAM and storage requirements.
OpenAI is moving in the same direction. GPT-6 Sol and Luna, launched on September 22, have API prices about 50% lower than the previous promotional price of GPT-5.6.
Anthropic launched Claude Opus 5.5 on the same day. Its input price per million tokens dropped from $5 for Opus 5 to $4, and its output price dropped from $25 to $20; for Cache Read, which is used very frequently in Agent scenarios, the price dropped from $0.5 to $0.2. Anthropic says that, when calculated together with token usage efficiency, the cost of running a typical task can fall by about 40%.
Viewed individually, each company's price cut can be explained as architecture optimization or economies of scale. Viewed together, the signal is clearer: AI intelligence itself is undergoing rapid price compression.
Open weights begin to capture volume, and tokens are becoming a "commodity"
Beyond falling prices, another change is taking place: more and more work does not necessarily require the most expensive model.
Here, it is necessary to first distinguish a concept that is often conflated. The show repeatedly uses "open source," but strictly speaking, open-weight is not completely equivalent to open-source. The former usually means the model weights can be downloaded and deployed independently, but the license, training data, or complete training code may not be fully open.
For enterprises, however, the commercial implication is very direct.
If an internal summary, customer service classification, code assistance, or routine Agent workflow can be completed well enough by a model that is ten times cheaper, then continuing to call the most expensive Frontier Model will become increasingly difficult to explain to the CFO.
Similar trends are already appearing in Vercel's AI Gateway data. Its Production Index released in September shows that among the tokens processed by the Vercel AI Gateway, the share of open-weight models has risen from 7% in December 2025 to 56% in August of this year, exceeding half for the first time.
At the same time, the average token cost has fallen to less than half of what it was five months ago, and in August alone, the unit token price dropped another 23.2%.
It should be emphasized here: this is production traffic data from the Vercel AI Gateway and cannot be directly extrapolated to mean that "56% of the world's AI tokens already come from open models." But it at least provides a sample from a real production environment: developers are beginning to migrate workloads to cheaper models more and more frequently.
Chamath summarized this change as "models beginning to cluster." His judgment is that more and more models have become close enough in capability across a large number of tasks to be substitutes for one another. So the question is no longer just "which model scores highest," but: how much more are you willing to pay for that last bit of performance gap?
Frontier models will not disappear, but they must prove "where they are expensive"
This is also the real business problem facing OpenAI and Anthropic.
David Sacks' view on the program is more optimistic than that of the other hosts. He believes that even if open-weight models take the vast majority of tokens, that does not mean Frontier Intelligence loses its commercial value.
The reason is that there will always be a set of tasks with an extremely high willingness to pay for "the best." For example, drug discovery, complex mathematics, cutting-edge engineering, cybersecurity, or highly competitive financial trading scenarios. For these tasks, if a stronger model can improve the success rate even slightly, the economic value it creates may far exceed the token cost.
Friedberg holds a similar view. The life sciences R&D organization he is part of uses Anthropic's models because, on some highly difficult scientific tasks, the differences in model capability are still large enough.
Therefore, what truly deserves attention is not "whether open models will eliminate Claude or GPT," but a more specific question: of OpenAI's and Anthropic's revenue, how much comes from genuinely irreplaceable frontier tasks, and how much comes from ordinary tasks that can ultimately migrate to cheap models?
If the vast majority of revenue comes from the former, model companies still possess strong pricing power. But if a large share of revenue exists only because enterprises are temporarily accustomed to using the strongest models, then as enterprises optimize AI costs, this portion of revenue may face increasing price pressure.
This is also why the customer structure and revenue quality of model companies may be more worth observing in the future than mere token growth.
From a capital markets perspective, this is especially important. What investors truly need to judge is not just "whether Claude or GPT can continue to grow," but how durable the frontier intelligence premium behind that growth really is.
As the "brain" becomes cheaper, value begins to shift to the "hands and feet"
Chamath offered another judgment on the program that is even more worth noting. If the model is seen as the "brain," what truly determines the final product experience is increasingly becoming the agent harness wrapped around the brain.
The so-called harness can be understood as the entire system that enables a model to actually "do things," including the browser, tool calling, memory, task planning, permission management, code execution, and the ability to connect with different software. The same model, placed in different harnesses, may ultimately perform completely differently.
This means that as foundation models gradually converge, the focus of competition among AI companies may begin to move up one layer.
The core question in the past was: who has the smartest model? Going forward, it may increasingly become: who can make this model actually complete work?
Meta's Muse is precisely a representative of this trend.
Meta released Muse on September 8, positioning it as a personal AI agent: it does not just answer questions, but can operate browsers and applications through an independent virtual computing environment to complete tasks such as sending emails and arranging travel.
On the program, Chamath mentioned that he used Muse to organize his personal inbox and also tried having it book flights and hotels. In his view, the most important significance of this type of product is not which model is used underneath, but that it packages originally complex AI capabilities into software that ordinary users can directly use.
This also explains why falling token prices do not necessarily mean the AI industry is losing value. On the contrary, cheaper intelligence may make more products economically viable. Thinner margins at the model layer do not mean the entire AI stack earns less. Value may simply be redistributed away from the models themselves to the companies that own users, workflows, and end products.
The next round of AI competition may no longer be about whose model is first
This shift is still in its early stages.
Capability gaps still exist between frontier models and open-weight models, and different benchmarks are far from representing all real enterprise tasks. What All-In calls "models are converging" is currently better understood as a business judgment rather than a completed industry fact.
But price changes are already clearer. Open-weight models are entering more and more production environments, API unit prices continue to fall, and closed-source model vendors are also actively cutting prices. At the same time, products such as Meta Muse are beginning to push AI from chat windows toward agents that can actually execute tasks.
What is truly worth watching next is no longer just how many more points the next generation of GPT or Claude scores on benchmarks.
More important are four variables:
First, whether the share of open-weight models in real production tokens can continue to rise; second, how quickly the cost per unit of intelligence can still fall; third, how many irreplaceable complex tasks high-priced frontier models can retain; fourth, whether agents can truly form stable usage frequency and commercial revenue.
If the first three trends continue to develop, AI models may increasingly resemble basic computing power in cloud computing: still important, but with prices constantly becoming more transparent and standards constantly becoming more unified.
And what truly determines a company's value in the next stage may no longer just be "who owns the smartest brain." Rather, it is who can turn increasingly cheap brains into products that people are genuinely willing to use and pay for.
Welcome to join the official BlockBeats community:
Telegram Subscription Group: https://t.me/theblockbeats
Telegram Discussion Group: https://t.me/BlockBeats_App
Official Twitter Account: https://twitter.com/BlockBeatsAsia
