Kimi K3 brings AI investors back to a core question: Does a cheaper and more efficient model mean a reduction in chip and memory requirements? Both Citigroup and Bank of America Securities lean towards a negative answer, suggesting that efficiency improvements may unleash more usage, thereby increasing overall resource consumption.
According to Windchaser Trading Desk, the Kimi K3 released by the Dark Side of the Moon has 280 trillion parameters, targeting a context window of 1 million tokens and long-cycle agent tasks. Citigroup semiconductor analyst Peter Lee believes that even if Kimi K3 is widely adopted, the general memory requirements such as server DDR5 and eSSD will continue to increase.
Bank of America Securities semiconductor analyst Vivek Arya also pointed out in a research report on July 17 that U.S. leading AI labs are more likely to increase computing power investment rather than reduce it. If China's open-source models continue to approach, leading players such as OpenAI, Anthropic, and Google will need larger training scales, heavier inferences, and faster product iterations to maintain differentiation.
This means that the market's pricing logic for Kimi K3 cannot solely focus on the cost reduction of single inference. More importantly, it is whether low-cost models will bring more invocations, longer task chains, and more token generation. If the answer is yes, GPUs, HBMs, DDR5, eSSDs, high-speed networks, and inference systems may continue to benefit.
Low Cost, High Performance, Triggering the 'Jevons Paradox'
Citigroup views Kimi K3 as a potential case of the 'Jevons Paradox' in the AI industry chain. The core of this paradox is that after a technological efficiency improvement reduces unit usage costs, usage may significantly increase, ultimately leading to a rise rather than a reduction in total resource consumption.
The allure of Kimi K3 lies in the combination of low cost and high performance. The publicly disclosed prices show that the input price per million tokens for cache hits is $0.3, and the output price per million tokens is $15. Its full weight plan will be disclosed on July 27, with the goal of supporting long-cycle agent work.
Citigroup's assessment is that a model price reduction will not automatically decrease hardware requirements. Instead, low costs may increase developers' and enterprises' willingness to invoke the model, driving more AI agent deployments. Agent tasks are not one-off queries but rather continuously generate, read, and process tokens in a chain of tasks. As invocation frequency and task length increase, the cost reduction per unit will be transformed back into total resource consumption.
Therefore, the impact of Kimi K3 on the semiconductor supply chain is not focused on whether "single inference is cheaper," but on whether the "total token count expands." This is the reason Citigroup is bullish on server DDR5 and eSSD demand.
Long Contextual Reasoning, Channeling Pressure to Memory
Kimi K3 adopts three key technologies: Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. Kimi Delta Attention is used to reduce the cost of a 1 million token context window, Attention Residuals are used to selectively retrieve representations between different depths of the model, and Stable LatentMoE enhances sparsity, activating only 16 out of 896 experts per token.
Technical documentation from the dark side of the moon states that this architecture enables Kimi K3's scaling efficiency to reach 2.5 times that of K2. High sparsity and long contextual capabilities are the key foundation for reducing operating costs.
However, Citigroup emphasizes that this does not mean the disappearance of inference-side resource pressure. Kimi K3 is still not a lightweight deployment solution and requires a multi-node cluster for operation, with super node configurations exceeding 64 GPUs. More importantly, long contextual and intelligent agent tasks will increase KV Cache usage, thereby adding to the memory burden on the inference side.
KV Cache requirements are directly related to server DDR5 and eSSD. DDR5 handles high-frequency data access, while eSSD benefits from larger caches and data storage requirements. For memory manufacturers, the key variables are not whether a single model is more cost-efficient, but how many times the model is invoked after the price reduction, how many tokens are generated per invocation, and how long the intelligent agent task chain is.
Model Convergence, Paradoxically Raising the Bar for Compute
Bank of America Securities' conclusions echo Citigroup's but with a focus on GPU and AI infrastructure. Vivek Arya believes that stronger Chinese open-source models may not necessarily lead U.S. AI giants to reduce their compute investment. Instead, narrowing the model gap will increase the cost for leaders to maintain their advantage.
Bank of America Securities points out that if open-source models continue to converge, OpenAI, Anthropic, and Google will need to rely on larger-scale training, more reinforcement learning and synthetic data loops, heavier testing-to-inference cycles, and faster product release cadence to maintain differentiation.
This logic does not depend on which model is leading in the short term. The big model leaderboard may rotate quickly, but what enterprises really purchase is stable, low-latency, highly available AI output, and a lower "per-unit of effective output" cost. As model capabilities converge, competitive pressure will percolate down to the underlying infrastructure, including GPUs, HBMs, high-speed networks, and inference systems.
Therefore, Bank of America Securities believes that the emergence of models like Kimi K3 should not be simply understood as "more efficient models, fewer chips." A more realistic path is that as model competition intensifies, frontrunners continue to ramp up computing power to maintain their lead.
MoE Architecture Shifts Bottleneck, Memory, and Interconnect to the Fore
Kimi K3 adopts the MoE architecture, which means the total parameter count and the actual parameters activated per inference can be separated. This change reduces some computational pressure but also shifts the infrastructure bottleneck from sheer computing power to memory access, expert routing, response latency, and interconnect capability.
Bank of America Securities emphasizes that MoE inference requires systems to more efficiently move data, schedule expert modules, and connect to larger compute clusters. NVIDIA suggests that modern MoE inference requires larger GPU domains. Its estimates show that the performance per watt of the GB300 NVL72 on leading open-source models can reach up to 25 times that of the Hopper platform.
CoreWeave's tests with Kimi K2.6 also indicate that even though open-source MoE models activate only a portion of the parameters each time, to maintain a lead in speed and cost-effectiveness, optimized NVIDIA GB300/GB200 NVL72 infrastructure is still required.
This implies that model weights may gradually become commoditized, but the GPUs, high-bandwidth memory, network interconnect, and inference systems needed to run the models will not lose value. Instead, as model invocation volumes increase, these aspects may become more central to the competition.
Token Usage Continues to Expand, Demand Side Has Not Peaked Yet
Bank of America Securities also quotes OpenRouter data indicating that model invocation volumes are still rapidly increasing. As a third-party API platform connecting multiple model labs, the usage of tokens on OpenRouter continues to grow, with the token usage of Chinese AI lab models surpassing non-Chinese lab models.

This data does not directly represent all enterprise and consumer scenarios, but it at least shows that developers' choices in an open model ecosystem are changing. As model prices decrease and capabilities improve, it may drive continued expansion of call volumes rather than being offset by the efficiency gains of a single model.
Enterprise paid usage is also on the rise. Ramp data shows that as of June 2026, approximately 55% of U.S. enterprises have paid subscriptions for AI models, platforms, or tools, surpassing the U.S. Census Bureau's BTOS survey estimate of 21%. By model, Anthropic has an enterprise adoption rate of 42.4%, while OpenAI is at 39.5%.
However, AI spending remains highly concentrated. The top 1% of enterprise users spend an average of about $4,833 per employee per month on AI, while the top 10% spend $516, with an overall median of only $11. The tech and media industry subscription rates reach 79.8%, with large enterprises at 65.5%, higher than mid-sized companies at 61.3% and small businesses at 48.7%.

This set of data supports one conclusion: AI usage is still in the process of diffusion. If low-cost models reduce the entry barrier, future incremental demand may come from more businesses, more developers, and more intelligent agent applications.
Shifting from "More Compute Efficient" to "More Workloads"
The joint conclusion of Citi and BofA Securities is that the significance of Kimi K3 lies not in whether a single model lowers the unit inference cost, but in whether the low cost unleashes a larger scale of workloads.
For Citi, the most direct beneficiaries are server DDR5 and eSSD. Long-context, KV Cache, and agent tasks will increase memory access and storage demands. For BofA Securities, what is more crucial are GPUs, HBM, high-speed networking, and inference systems, as model competition will compel front-runners to keep investing in infrastructure.
BofA Securities also notes that Kimi K3 involving "45nm open-source EDA" should not be simply seen as commercial EDA being replaced. Instead, this indicates that chip design still cannot do without EDA tools. In advanced processes, commercial EDA vendors like Cadence and Synopsys still hold key positions.
The risk lies in another scenario: if model compression, inference optimization, and hardware efficiency improvements outpace new workload additions in the long run, the expansion of AI infrastructure may undergo a phased cooling-off. However, in the current frameworks of these two investment banks, Kimi K3 appears more like a catalyst to boost usage rather than a signal of waning chip demand. After efficiency gains, what the market needs to focus on is more tokens, more inferences, and more underlying hardware consumption.
Welcome to join the official BlockBeats community:
Telegram Subscription Group: https://t.me/theblockbeats
Telegram Discussion Group: https://t.me/BlockBeats_App
Official Twitter Account: https://twitter.com/BlockBeatsAsia
