Morgan Stanley Estimates CapEx Profits: From Renting GPUs to Selling Tokens, How Does AI Infrastructure Earn a 25%-50% Return?

Bitsfull2026/07/29 16:0013890

概要:

Morgan Stanley Carves Out Three Exit Routes


In a report dated July 27, Morgan Stanley provided a rough estimate of AI capital expenditure payback: If GPU utilization, rental costs, token throughput, and API pricing meet the baseline assumptions, the Generative AI infrastructure and model API business have the potential to achieve approximately a 25%-50% capital return.


This directly addresses the market's biggest concern about large tech companies. Microsoft, Amazon, Google, and Meta continue to invest heavily in GPUs and data centers, with capital expenditures increasing. Investors are eager to know whether this money will translate into profits or merely inflate depreciation, energy consumption, and R&D expenses.


This assessment breaks down the monetization of GenAI into three paths: renting out GPU compute as IaaS, providing Model APIs through proprietary infrastructure, and renting third-party compute to offer Model APIs. The first two paths are more suitable for large tech platforms with data centers, customer access points, and product distribution capabilities, while the third path is more susceptible to GPU rental price pressures.




$1.4 Trillion in Spending, More Than Just GPU Purchases


The backdrop of this calculation is that AI data center construction is entering a phase of heavier capital expenditure. Publicly quoted figures suggest that Morgan Stanley expects hyperscalers' capital expenditure to exceed $1.4 trillion by 2028, with computing capacity potentially doubling from 2025 to 2028 to around 120GW.


All this capacity will not immediately translate into revenue. Top model training, model maintenance, and model iteration at different scales will still consume a significant amount of compute power. Training itself does not directly generate revenue but incurs depreciation, energy, and operational costs.


What truly affects the return on investment is whether the remaining compute can be fully absorbed by inference, APIs, enterprise software, and cloud services. In more direct terms, how many hours a GPU can be sold in a year, how much revenue can be generated per hour, how many tokens can be processed per second, how much can be earned per million tokens — these factors will determine whether AI capital spending can achieve a cash return.


This is also the most newsworthy aspect of this report. In the past, the market has seen more of an upward revision of Capex figures, but now Morgan Stanley has provided a set of unit economics calculations: under the baseline assumption, AI infrastructure can withstand high depreciation pressure and achieve a return level close to high-quality cloud or software businesses.


Rent GPU: At 75% Utilization, IaaS Return is about 31%


The first path is for major cloud providers to offer GPU computing power as Infrastructure as a Service (IaaS). The baseline scenario assumes that 1GW of capacity corresponds to approximately 410,000 NVIDIA GB300 GPUs, with a 75% utilization rate and an hourly rental rate of $8.5.


Under these assumptions, the GPU rental business generates approximately $22.9 billion in revenue per GW-year, with an incremental EBIT margin of about 67% and a capital return of around 31%. As long as the GPU supply can be filled by sufficiently high demand, IaaS is not just about low-margin hardware rental but is closer to high-utilization infrastructure business.




The advantage here comes from the major cloud providers' existing power, data centers, customer relationships, and cloud platform sales capabilities. If the additional GPU capacity aligns with existing customer demand, revenue conversion will be more direct.


However, this path is very sensitive to price and utilization rates. A decrease in GPU rental rates, underutilization, or rising energy costs will all depress returns. With the entry of new players and the launch of ASICs and next-generation GPUs, the unit computing power price may also decline, making it uncertain whether the $8.5/hour rental rate can be sustained in the long term.


Model API Is More Profitable, Betting on Token Consumption


The second path is to provide Model API using proprietary infrastructure. The baseline scenario assumes that 65% of the capacity is used for inference, with each GPU capable of processing approximately 2,750 tokens per second, and a blended pricing of $1.75 per million tokens.


In this scenario, the Model API's incremental EBIT margin is about 75%, with a capital return of over 40%. Across different reprints, this return is about 40%-46%, but the trend is consistent: adding proprietary computing power to model services yields higher returns than just renting out GPUs by the hour.


The reason is not complicated. Cloud providers and model providers sell not individual GPU hours, but model capabilities, inference services, and API calls. Revenue is tied to token consumption, providing a larger profit margin.




This also explains why large tech companies are buying chips on one hand while embedding AI capabilities into search, office software, ads, e-commerce recommendations, and developer tools on the other. What they truly aim to do is not just lease out idle compute power, but turn inference into a high-frequency, billable, embeddable revenue stream within existing products.


Risks are also concentrated at the token level. Token throughput depends not only on the chip but also on the model's parameter size, software stack, I/O ratio, and inference optimization ability. If the model is more efficient, a single GPU can process more tokens, increasing ROI. If open-source models and price wars drive down revenue per million tokens, the profit margin will be diluted.


Renting Third-Party Compute Power for APIs Eats Into Profits More Easily


The third path is for model companies to rent third-party infrastructure and then provide Model APIs to customers. In a baseline scenario, GPU rental is about $7.75 per hour, with an incremental EBIT margin of about 31%, resulting in an after-tax ROI or profit margin of around 25%.


This path can still be profitable but not as lucrative as owning a self-owned infrastructure platform. The party renting GPUs must first pay for compute power rental and then bear the costs of model training, inference, services, and sales, leaving a narrower profit margin.




This is also why Microsoft, Amazon, Google, and Meta have a relative advantage. They have the capital, data centers, customer access, and product distribution capabilities to choose the most suitable monetization method between IaaS and APIs. Pure model companies or intermediate layer applications are more easily squeezed by compute costs if they lack strong pricing power.


This does not mean third-party model companies have no opportunity. High-quality models, vertical scenarios, enterprise customization, and application-layer distribution may still form differentiators. However, from a capital return perspective, the party with low-cost computing power and high utilization infrastructure is more likely to retain profits.


The Landing Point is Four Large Tech Platforms, But High Returns Have Not Been Realized Yet


Morgan Stanley continues to favor Microsoft, Amazon, Meta, and Google in this framework. Their common advantage is that they can both expand AI infrastructure and have ready-made products and customer entry points to meet inference demands.


Microsoft is rated Overweight with a target price of $600. Alphabet, Google's parent company, has a target price of $400, Meta is listed as a Top Pick with a target price of $775. Amazon is also included in the direction benefiting from large AI platforms, but its specific target price varies in public sources, making it inappropriate to use a single number as a conclusion.


The landing points of these stocks should not be understood as AI capital expenditures having already paid off. The short-term financial reports of the four companies may still face pressures from depreciation, energy consumption, chip upgrades, and R&D expenses. Whether AI investment can translate into shareholder returns ultimately depends on whether the revenue growth from inference can outpace cost expansion.


A return rate of 25%-50% is also just a scenario calculation, not a realized financial result. A 75% utilization rate is a high threshold, and if enterprise AI applications land slower than expected, idle GPUs will directly drag down returns. Token pricing is also volatile, and open-source models, more efficient smaller models, and cloud provider price wars could all reduce income per million tokens.


The report provides more of a yardstick: as long as large tech companies can maintain AI infrastructure at high utilization and turn token consumption into revenue through Model APIs, cloud services, and existing products, AI capital expenditures will not just be a cost sinkhole. Conversely, if rental rates, utilization, and inference adoption rates fall short, a 25%-50% return will quickly become stuck in the model spreadsheet.



Welcome to join the official BlockBeats community:

Telegram Subscription Group: https://t.me/theblockbeats

Telegram Discussion Group: https://t.me/BlockBeats_App

Official Twitter Account: https://twitter.com/BlockBeatsAsia