The Silicon Valley That Copies China's Homework Is Now Doing AI Applications

Bitsfull2026/08/04 11:3619315

Summary:

「Invented in the U.S., Scaled in China」, an exception has emerged.


Silicon Valley has started copying China's homework.


Recently, the U.S. has produced a batch of low-cost open-source models aimed at catching up with China's open-source models in terms of capabilities while also pricing them at the same level.


Mira Murati's Thinking Machines Lab released the first model, Inkling, last month. The official statement mentioned two key points: the underlying architecture was referenced from DeepSeek-V3, and the training data was generated by Kimi K2.5.


Some of the most expensive people in Silicon Valley are now using the framework of Chinese open-source models and feeding them with data generated by Chinese models. Just two years ago, this scenario would have been unthinkable.


Earlier this year, Arcee AI's Trinity-Large-Thinking scored 91.9 on PinchBench, 1.4 points lower than Anthropic's Opus 4.6, but priced at only 4% of the latter.


Arcee AI's CEO, Mark McQuade, said, "We are working hard to help the U.S. catch up with and surpass China."


Following the model, Silicon Valley is now not only copying the models but also the prices set by Chinese companies.


Architectures can be emulated, training methods can be reproduced, and data can be handed over to other models for mass generation. Once the weights are released, developers can continue the development. As for pricing, it's a simpler matter – as long as one company slashes prices, the rest will sooner or later follow suit.


Over the past three years, large model companies have built a towering building based on parameters, computing power, and closed-source systems, hosting grand banquets. Now, this building is on the verge of collapse.


Sixteen Days


In the early hours of July 17th, the dark side of the moon released Kimi K3. With a total of 28 trillion parameters and 1 million tokens in context, it is the world's largest open model.


In the Frontend Code Arena programming leaderboard, K3 scored 1679, surpassing Claude Fable 5 and GPT-5.6 Sol. This is the first time an open model has outperformed closed-source models. In the seven frontend sub-directions, it secured the first position in six.


Musk left a comment below the assessment results saying, "Impressive."



The API of K3 is not cheap, charging $100 per million output tokens. However, in the SuperCLUE's horizontal testing, its task score was 18% higher than the second place, and after running the same batch of tasks, the average cost actually decreased by 16%.


Whether a model is expensive or not should not be judged by the price per million tokens, but by how much it costs to complete a task from start to finish. The farther these two numbers are apart, the more meaningless the price tag itself becomes.


In the late night of two days later, the dark side of the moon paused the C-end new user subscription. The 48-hour request volume far exceeded the estimate, and the computing power was insufficient.


On July 27, K3 opened up its full weight, along with the technical report and supporting infrastructure.


Next up was DeepSeek.


The market had been waiting for the official release of V4. At the end of April, DeepSeek first released two preview versions, V4-Pro and V4-Flash, both supporting a context of 1 million tokens. The official term for this was "the era of million-token context for all." There was little news for the next two to three months until August 2, when V4 Flash was officially released with open weight. The total parameter was 284 billion, with only 13 billion activated in a single inference, and the score of DeepSWE programming test rose from 7.3 to 54.4.


What's even more aggressive is the pricing. In May, V4-Pro was directly reduced to one-fourth of the original price, marking the fourth price adjustment by DeepSeek in a month. When Flash was officially launched, everyone was still surprised by its cost-effectiveness.


"You get what you pay for" is a thing of the past. Now, Chinese models are becoming more competitive, charging less and less. Foreign competitors are naturally uncomfortable, but there's nothing they can do. Users have seen affordable and user-friendly products and no one is willing to be a sucker again.


Subsequently, Alibaba announced that it will release the weight of Qwen3.8-Max next week. The Max series, which has always been in the flagship closed-source position, has also started to open up. Its international pricing is approximately 40% of Opus 5 in, and the output is only 24%.


All these events took place within a mere 16 days.


In the past, model companies relied on scarcity for pricing; only I could do it, so you had to pay my price. On the leaderboard, the difference in price could be tens of times greater on just a few points difference.


The technical roadmaps of the three companies, Kimi, DeepSeek, and Qwen, are different, but they are all doing the same thing—elevating capabilities, opening up weights, lowering prices, and swiftly incorporating them into products.


Wanting to Copy Homework, But No One Is Paying


Some in the U.S. have also understood this trend.


By the end of 2025, Arcee AI almost entirely bet the company's money. In 33 days, using 2048 B300 chips, $20 million, and with only 30 employees, they created the Trinity Large with a total parameter of 400 billion, a single inference activation of 13 billion, and nearly half of the training data being synthetic. In July of this year, Arcee also signed a partnership with the U.S. Department of Energy, preparing to develop a model for scientific research.


The team, achievements, costs, and government orders, everything is in place. Yet, the door to financing remains firmly shut.


Arcee has raised a total of $50 million to date, with a valuation of $240 million. In today's era of large models, this amount of money wouldn't even cover the cost of the table, let alone grant access to the VIP room.


CEO Mark McQuade said, "Almost all top-tier VCs have rejected us."


The American open model has not yet directly competed with China and has been intercepted by its own investment committee.



In the first quarter of 2026, global AI startups raised $255.5 billion, with close to two-thirds concentrated in three deals: $122 billion for OpenAI, $30 billion for Anthropic, and $20 billion for xAI. The money has mostly flowed into closed-source companies.


Joe Floyd, an early investor in Arcee and a partner at Emergence Capital, said that the reasons for rejection he heard from many VCs were similar:


"I don't want this model to succeed. I don't want to invest because it will harm my investment in Anthropic and OpenAI."


Ironically, those willing to spend money on the open model are chip sellers.


In March of this year, NVIDIA disclosed in SEC filings that it is prepared to invest $26 billion over the next five years to support open weights. Reflection AI, Poolside, and Thinking Machines Lab have all received its funds.


This account's Huang Renxun understands it better than anyone else. The proprietary model makes money from API, while the open model burns through GPUs. The cheaper the model is sold, the more developers will kick up a fuss, and NVIDIA's GPUs will sell better.


On the evening of July 24, Huang Renxun registered an X account and tweeted for the first time by retweeting an open letter titled "Open Weights and America's Leadership in AI.” The letter mentioned that whether the U.S. can continue to lead this industry should not only depend on whether they have the most advanced model but also on whether they have an ecosystem that can spread across various industries.


Later, founders of nearly 200 U.S. startups jointly wrote to the Trump administration, opposing the ban on Chinese open models. After reading this letter, the most interesting part was that it didn't contain anything significant. They didn't speak positively about China, nor did they advocate for openness. Their only argument was that the cheap Chinese models had already become production tools, and if they were to be banned, they would have to go back and buy expensive American APIs.


McQuade also mentioned another sentence during an interview, saying that mass adoption is only a matter of time.


Demand Driven by Affordability


While Silicon Valley is still debating whether to proceed with open weights, China has already taken a step forward.


The phrase "Year of AI Applications" has been repeated many times over the past three years. However, to determine whether a technology has truly entered the application stage, one only needs to look at two things: whether users are willing to pay for it and whether enterprises have integrated it into their operations.


This year, the annual revenue of Claude Code and Cursor has both exceeded $1 billion, and the penetration rate of AI programming tools among individual developers is close to 50%.


Enterprises have undergone even greater changes. According to a study by the Sand Dune think tank, the adoption rate of AI Agents in Chinese enterprises has increased from 17.3% at the end of 2024 to 40.3% in mid-2026.


For this, Li Yanhong proposed a new indicator called DAA, the Daily Active Agents count. Previously, internet companies counted how many people opened their products every day. From now on, they must count how many Agents are working for people every day, and how many have actually delivered results.


The surge in usage is even more astonishing. According to CCTV's calculation method, China's daily token call volume increased from about 100 billion in early 2024 to 140 trillion by the end of March this year, a growth of over 1,000 times in just over two years.


While models are becoming increasingly affordable, computing power is becoming scarcer. In March this year, Tencent Cloud, Alibaba Cloud, and Baidu Cloud successively raised their computing power prices by around 30% within ten days. In the first half of the year, as the Intelligent Agents surged, the leasing price of inference computing power increased by over 40%.


The reason is that the way people use AI has changed. Previously, a question-and-answer interaction didn't cost much in terms of tokens. Now, for an Agent to truly complete a task, they need to have the context, use several tools, double-check back and forth, and if they make a mistake, they have to start over. Though the unit price has decreased, the consumption has multiplied.


The price reduction did not shrink the market; instead, it brought in all the previously unquantifiable demands.



In 1981, IBM opened up the PC architecture, and compatible machines quickly flooded the market, driving hardware prices down. In the end, those who really made big money were not the ones making the machines but the ones with software like Lotus 1-2-3, WordPerfect, and later Microsoft.


Today, this path is being followed almost identically. While many are focused on the slightly lacking underlying capabilities, companies above have already started to snatch users, entry points, and workflows.


A U.S. company can study DeepSeek's architecture, use data generated by Kimi, or even spend $20 million to train a decent open model. But it's not so easy to directly apply it.


The application doesn't have pre-downloadable weights or benchmarks. Whether it can run or not depends on the tricks of the trade accumulated by a company over more than a decade.


In the past, major companies tended to turn models into standalone products, waiting for users to specifically open a chatbox. Now, models are starting to integrate into DingTalk, Feishu, and WeCom, directly becoming part of the original organization and workflow.


In the U.S., an office Agent might grow along Gmail, Slack, and Salesforce. In China, it might first enter DingTalk, Feishu, and WeCom, then move on to financial software, supply chain, factory systems, and government platforms. Even if both sides use the same model, the final products will be different.


After the Price Reduction


From 2022 to 2025, the U.S. increased its chip restrictions on China. H100 is not for sale, B300 is not for sale, and advanced processes and manufacturing equipment are also restricted. Under such conditions, Chinese companies have created the world's largest open model.


By 2026, the U.S. began discussing restricting the weights of Chinese models. OpenAI and Anthropic have reportedly expressed concerns to regulators, Congress is also studying whether model weights can continue to flow across borders.


The knife hasn't even fallen, but cries of pain are already heard among allies.


The chip ban targets the supply chain of a Chinese company, but the first cut of the model ban hits the cost of an American startup. The joint letter made it very clear that banning will not stop the spread; it will only force developers to hurry up and download all available models before the door closes.


A chip is a physical entity that can block orders, logistics, and factory manufacturing. Once the weight is released, it can be endlessly downloaded, mirrored, quantized, fine-tuned, and distilled.


What's even harder to control is the price.


The U.S. still leads in the most cutting-edge closed-source models. OpenAI and Anthropic are still at the forefront, and training the top-tier models still heavily relies on NVIDIA's GPUs, which won't change in the short term.


However, when it comes to application, the outcome is rarely determined by a single top model. The competition is about whether the model is affordable enough, if enterprises have a ready-made digital entry point, and if developers can truly integrate into the business processes.


Looking back over the past thirty years, China has won many times, but the winning strategy has been largely the same.


Mobile payments, e-commerce, food delivery, live streaming—all have involved taking technologies invented elsewhere and integrating them into the daily lives of over a billion people, turning them into massively scalable businesses. The underlying foundation has long been provided by others. The transistor came from Bell Labs, x86 belongs to Intel, TCP/IP was funded by the U.S. government, and Windows, Android, and iOS all grew up on the West Coast. Even the most commonly used TensorFlow and PyTorch in the era of deep learning originated from Google and Meta, respectively.


"Invention in America, scale in China"—this statement has been made for thirty years, with very few true exceptions.


Now, an exception has emerged.


The technical specification from Thinking Machines Lab is very clear. One of Silicon Valley's most prominent and wealthiest new companies referenced the architecture of a Chinese company at the model's base layer and later trained using data generated by another Chinese model.


In the past, Silicon Valley procured capacity, recruited talent, or sought markets from China. Now, it is starting to directly follow the technology path validated by Chinese companies.


The Silicon Valley worship certainly has a history. For the past few decades, most of the essential underlying technologies of the computer industry have indeed originated from there. From the chips inside the machines to the IDEs that programmers open every day, Silicon Valley has long been positioned upstream in the technology chain.


But in the tech world, what matters in the end is the outcome. Whoever can propose a new architecture, bring down costs, and make peers start to imitate, is eligible to redefine the upstream.


Now, the names of Chinese companies are beginning to appear on this list.


Silicon Valley will most likely catch up in the future, as there are plenty of engineers, capital, and chips there. But applications that have already been integrated into real business processes will not pause and wait for others. Whoever dives into the enterprise's workflow first will have the opportunity to gain customers, data, and the next round of improvements. The starting line has never been the same.


The era of AI applications will not choose an auspicious day to open with great fanfare. Its initial appearance is in financial statements, in skyrocketing API calls, and in a programmer once again depleting the quota in the middle of the night.


Mass adoption is only a matter of time. This statement was made by an American.



Welcome to join the official BlockBeats community:

Telegram Subscription Group: https://t.me/theblockbeats

Telegram Discussion Group: https://t.me/BlockBeats_App

Official Twitter Account: https://twitter.com/BlockBeatsAsia