J.P. Morgan Insights: Token and H100 Synchronized Price Drop, Is AI Cost Cooling Down?

Bitsfull2026/07/28 16:3119761

Summary:

The Cost of AI is first cooling down from the inference and leasing side, with the AI chip crunch expected to ease by 2026


JPMorgan's July edition of "Data Center Watch" revealed that AI inference call prices and high-end GPU short-term lease prices cooled off simultaneously at that time.


The direct implication of this development is that the most tense phase of AI infrastructure is starting to relax. For regular users and enterprise clients, the decrease in Input Token prices will make calling large models cheaper. For AI application companies and cloud providers, the drop in H100 rent indicates that the computing power market is no longer as queue-ridden as it was during the previous scarcity period.


However, this is not a signal of waning demand. This tracking sample shows that LLM Token processing volume has increased several times since the beginning of the year, with top model utilization still in the 70%-80% range, and high-performance GPU utilization exceeding 80%. The cooling of prices is more like a part of supply catching up with demand with open-source models and inference efficiency, rather than a sudden weakening of AI demand.


Token Prices Dip Slightly, Volume Still Amplifying


In July 2026, the input Token price of mainstream models continued to dip. According to the report sample, the price per million input Tokens of some mainstream models generally ranged from $0.2 to $5, with an average monthly decline of about 3% in July.



Some context is needed for this range. The output Token price of high-performance closed-source models such as GPT-4o, Claude 3.5 Sonnet is usually significantly higher than the Input Token price, and prices for lightweight models like GPT-4o mini cannot be directly compared with flagship models. This means that the lower prices in the report more reflect the input Tokens, lightweight models, or hosted provider sample calibers and cannot be understood as the complete call cost of all mainstream models has dropped to the same range.


Even so, the direction is still clear. For AI application companies, the unit inference cost of the same question-answering, code generation, search enhancement, or customer service calls is decreasing. Over the past few years, one of the most difficult accounts to settle in the commercialization of large models is "the more you use, the more you lose," and the downward trend in Input Token unit prices has at least brought some high-frequency applications closer to the affordable cost range.


More importantly, the price reduction occurred amidst a surge in inference volume. From the first half of 2026 to July, the processing volume of LLM Tokens in the sample report increased several times compared to the beginning of the year, with some subcategories experiencing growth of over fourfold. Leading models such as OpenAI, Anthropic, Meta, among others, maintained a steady high utilization rate of 70%-80%, while the widespread deployment of Llama series open weight models also significantly contributed to the increase in inference volume.


What the market witnessed was not a price reduction due to lack of usage but rather a unit price reduction driven by a higher volume of inferences, as model service providers, the open weight ecosystem, and infrastructure suppliers collectively pushed the price per unit down.


Open Weight Models No Longer Solely Relying on Price Competition


The rapid growth of open weight models was another major theme in the July Token market.


The report indicates that in July, the usage of open weight models increased by 26% month-on-month and surged by 63 times year-on-year compared to the same period last year; the volume-weighted average price rose by 36% month-on-month, and Token expenditure increased by 71%.


This suggests that open weight models are not merely squeezing closed-source models through low prices. With performance gradually approaching cutting-edge levels, the call prices, usage, and expenditure of some open weight models are increasing concurrently.


In July, the expenditure share of open weight models rose to 20%, up from 12% in June and 10% in May. Although closed-source models only represent 29% of the Tokens' total supply, they still contribute 80% of the expenditure, indicating that high-performance closed-source models still capture the majority of the commercial value, but open weight models are quickly catching up.


By Token usage, the top five models in July were MiMo v2.5, DeepSeek v4 Flash, GLM 5.2, DeepSeek v4 Pro, and MiniMax M3, collectively accounting for 50% of the Tokens' total supply.


By Token expenditure, the top five models were Claude Opus 4.8, Claude Opus 4.7, Fable 5, Kimi K3, and GPT-5.6 Sol, comprising 54% of the total expenditure. With Kimi K3 entering the top five expenditures, it signifies that open weight models are moving into the high-value territory traditionally dominated by closed-source models.



For enterprise clients, an expanded choice of models helps reduce reliance on a single closed-source vendor. However, as seen in this report, the changes brought about by open weight models go beyond price reduction and also involve competing for higher-value inference scenarios and budgets.


H100 Rental Price Drops, But High-End Cards Not in Oversupply


The GPU rental market has also experienced a similar structural differentiation.


In July 2026, the average rental price of H100 in the non-hyperscale cloud provider market was $2.68 per GPU hour, a 1.1% month-on-month decrease. This is the first month-on-month drop in H100 rental prices after seven consecutive months of increase.


This decline is much lower than the 15% to 25% range and cannot be described as a "20% drop in H100 rental prices." The report still tracks the average monthly rental price calculated per GPU hour, rather than weekly rental prices per card.


This indicates a supply-side response. More GPUs are entering the cloud rental market, intensifying the competition among cloud service providers and computational power platforms. Short-term prices are no longer solely determined by extreme scarcity. Compared to the high-end GPU scarcity phase, the pressure of unavailability, both for purchase and rental, as well as queuing for cards, has eased.


However, high-end resources are still not cheap. The H100 price remains significantly higher than the A100, with GPU overall utilization rates maintained at 75%-90%, and high-performance GPU utilization rates exceeding 80%. As long as the utilization rate remains in this range, the decline in rent appears more like squeezing out a portion of the scarcity premium, rather than a comprehensive oversupply of computing power.



For AI companies, this change brings two layers of impact.


First, the cost of inference business is more likely to decrease. Inference requires a continuous, stable, low-latency supply of computing power. The decline in GPU rental prices will improve the cost structure of API service providers, AI search, code assistants, and enterprise AI products.


Second, the cost pressure of training large models is not easy to alleviate simultaneously. High-end training relies not only on the number of GPUs but also on cluster interconnection, memory bandwidth, scheduling efficiency, and stable power supply. The drop in H100 rental prices does not mean a proportional decrease in large-scale training costs, nor does it mean that all AI startups can access cluster resources of equal quality at a low price.


Memory Price Hike Slows Down, High-End Cluster Price Reduction Not as Direct as Short-Term Rental


The least thorough loosening of prices is seen in memory.


In May to June 2026, the sample data in the report indicates a rapid increase in DRAM spot prices, with categories like DDR5 16Gb experiencing a cumulative increase of over 90%-100%, stabilizing only in July. In contrast, HBM, due to AI accelerator demand, remains at a high level, with contract prices lagging behind spot prices by 1-2 quarters.


For high-end AI accelerators, the GPU chip itself is not the only bottleneck; HBM, advanced packaging, and data center delivery will also impact real-world supply. Even as some GPU prices in the rental market fall, upstream memory and high-end cluster delivery will continue to limit the overall cost reduction rate.



This is also why the "tightest moment of computing power begins to loosen" cannot be directly translated as "end of computing power shortage." The decline in rental market prices indicates that some short-term supply pressures have eased. However, high-end training and large-scale inference clusters are still constrained by memory bandwidth, advanced packaging, and the pace of next-generation GPU shipments.


Time calibration is essential. The report reflects price and utilization changes in a July 2026 sample, making it more suitable as a snapshot of the easing AI computing power in mid-2026, rather than a generalization of prices for all cloud providers, regions, and long-term contracts.


Cost Reduction, Model Vendor Profit Pressure Shift


The decrease in Token and GPU rental prices is a positive development for AI application companies, but it may not be entirely good news for model service providers.


Lower calling prices help stimulate demand, increase API usage, and enable more enterprises to integrate large models into real business processes. The issue is that if prices drop faster than the efficiency of inference improves, model vendors and cloud platforms' profit margins will be squeezed. Particularly in the context of open-weight model catching up and enhanced customer bargaining power, the high-price space of closed-source APIs will continue to be challenged.


For hardware and cloud service providers, the drop in H100 short-term rental prices also signifies the normalization of the ultra-high returns brought about by extreme shortages in the past. Variations in long-term contracts, regional differences, batch scale, and spot price calibers will all lead to significant differences in the actual prices obtained by different customers. By just looking at the weekly rental averages, one cannot directly infer that all cloud providers' revenues will decline simultaneously.


The most explicit signal from this price tracking is that by mid-2026, part of the AI infrastructure supply has caught up with demand, and the unit calling and short-term costs are beginning to cool down. However, the boundaries are also clear: Token usage is still rapidly increasing, high-end GPU utilization remains high, output Token prices are higher, and HBM remains tight. Cost reductions are occurring first in the segments of the market that are easier to commoditize, with the most challenging areas to loosen being high-end clusters and upstream memory supply.



Welcome to join the official BlockBeats community:

Telegram Subscription Group: https://t.me/theblockbeats

Telegram Discussion Group: https://t.me/BlockBeats_App

Official Twitter Account: https://twitter.com/BlockBeatsAsia