Open weights, model routing, and architectural efficiency improvements will reduce the cost per unit of intelligence but may amplify total compute demand through Jevons’ paradox. As the cost of simple API calls declines, enterprises will deploy more continuously running agents, parallel sub-agents, long-context analytics, code automation, and real-time multimodal services.
Zhitong Finance APP has learned that the recent evolution path of large AI models and their market penetration trends have stunned institutional investors focused on the valuation prospects of closed-source AI application leaders such as OpenAI and Anthropic. Notably, global tech industry leaders—including NVIDIA (NVDA.US), Microsoft (MSFT.US), Meta Platforms (META.US), the U.S. tech veteran IBM (IBM.US), and Silicon Valley venture capital powerhouse Andreessen Horowitz—are urging the U.S. government to support the development trajectory of open-weight large AI models. Over 20 of the world’s leading technology companies, including NVIDIA and Microsoft, signed a joint open letter endorsing open-weight AI models. Although OpenAI and Anthropic did not sign the letter, prominent figures such as Elon Musk, the world’s richest person and founder of SpaceX, and Microsoft CEO Satya Nadella have publicly voiced their support.
In the AI developer ecosystem, 'open-weight AI' typically refers to large AI models whose trained parameters can be downloaded by developers for deployment, fine-tuning, and inference on their own servers or cloud environments. However, the model developers may not necessarily disclose training data, data processing methods, training code, or full architectural details, and may impose license restrictions on commercial use or specific applications.
Strictly speaking, open-source AI encompasses a broader scope: in addition to model weights, it should also provide sufficient code, architecture, and training data information to enable users to study, modify, and reproduce the system, while guaranteeing fundamental freedoms such as use, study, modification, and redistribution. Thus, open-weight AI means the 'final model product can be used and adapted,' whereas truly open-source AI further discloses 'how the model was built.' The former can be part of the latter, but open-weight does not necessarily equate to open-sourcing all technical components and code.
It is understood that this significant open letter from the AI technology sector, released last Friday, comes at a time when the White House, citing national security concerns, is considering whether to ban Chinese open-weight models. More than 20 companies signed the letter, including A16z, Dell, Microsoft, Meta, NVIDIA, and Palantir—leaders across the AI application software landscape.
For global investors focused on the AI computing power supply chain and investment trends in the AI super bull market, open-weight models, improved model routing, and architectural efficiency will reduce the per-unit cost of intelligence but could amplify total compute demand through the Jevons paradox. As the cost of simple API calls declines, enterprises are likely to deploy more continuously running agents, parallel sub-agents, long-context analysis, automated coding, and real-time multimodal services. While individual tasks may consume less compute, the number of tasks, inference steps, and deployment nodes could grow even faster. Even if ultra-large MoE models like KIMI K3 activate only a small subset of experts, they still require massive weights to be distributed across high-capacity memory and high-speed interconnect clusters. Consequently, the proliferation of open models will diffuse compute demand beyond a few closed-source labs to cloud service providers, sovereign clouds, enterprise-scale AI data centers, and on-premises inference clusters.
The recently much-discussed Jevons Effect—also known as the Jevons Paradox—is a counterintuitive economic theory stating that when technological advances improve the efficiency of using a particular resource (such as energy, raw materials, or AI computing infrastructure), they lower the unit cost, thereby stimulating a significant expansion in market demand and ultimately leading to a net increase—not decrease—in total resource consumption. The concept was first introduced by British economist William Stanley Jevons in his 1865 book, The Coal Question.
Why have U.S. tech giants suddenly begun supporting open-weight AI?
Microsoft and other signatories stated in the open letter: 'Open-weight large AI models expand opportunities for businesses worldwide to participate in the economic prosperity of the AI era.' The letter emphasized that organizations can build upon advanced AI models without having to train models from scratch or pay the high token costs associated with cutting-edge proprietary models.
The tech giant added that open-weight models foster competition and ensure that the benefits of AI are more widely shared rather than concentrated among a few players.
‘Open-weight enables every organization to match the right model to the right task at an appropriate cost—reserving frontier-scale capabilities for truly frontier problems, while integrating, running, and deploying efficient, specialized models across all other scenarios.’
It is understood that NVIDIA CEO Jensen Huang has been a vocal advocate for open-weight AI. When sharing the joint letter, he wrote: ‘I’m sharing a letter co-signed by NVIDIA explaining why open large AI models are crucial. AI will ultimately transform every industry, empower every business, and be built and led by every nation.’
Unlike closed, proprietary AI systems dominated by companies such as OpenAI and Anthropic, open-weight models allow developers and enterprises to download nearly all trained parameters—commonly referred to as 'large model weights'—and run them on their own custom-built AI computing infrastructure. In essence, this concept mirrors the open-source software movement that transformed the global computer industry decades ago.
Open-weight models differ from fully open-source AI. Open-weight models typically release the trained model weights for free public download and enable one-click deployment, but do not necessarily disclose the underlying training data or all source code related to the large AI model.
Training advanced AI models can cost billions of dollars. The open letter argues that open-weight models can largely address this challenge, as developers can build upon or fine-tune existing AI models rather than creating one from scratch. For smaller enterprises, this significantly lowers the barrier to entry in the race to develop AI-era products.
According to the letter, open-weight models not only foster competition among AI developers but also promote healthy and vigorous competition among cloud service providers, chipmakers, software vendors, and AI application developers.
“This competition can spur innovation, reduce costs, and broadly distribute the economic benefits of AI,” as open-weight AI substantially lowers a range of costs associated with developing AI products.
In contrast, developers of closed AI models—such as OpenAI, Anthropic, and Alphabet’s Google (GOOGL.US)—clearly have strong economic incentives to maintain exclusive control over their most advanced technologies.
How do tech giants like NVIDIA view the security risks associated with open-weight large AI models? The open letter adds that cybersecurity defense efforts require priority access to these state-of-the-art AI capabilities—whether open- or closed-source—to counter increasingly sophisticated cyberattacks in the AI era. Researchers argue that open-weight large AI models enable more researchers to identify vulnerabilities, improve security systems, and conduct large-scale, independent system evaluations without relying solely on the original model developers.
Kimi triggers a disruptive wave combining 'open weights + low-cost tokens,' with model routing taking over as the AI gateway.
The real disruption brought by Kimi lies not merely in 'lower pricing,' but in the combination of high performance-to-cost ratio, open weights, long context windows, and API compatibility—which collectively lower the barrier to model migration. Kimi K2.6 is officially priced at approximately USD 0.95 and USD 4 per million input and output tokens, respectively, with cached inputs costing just USD 0.16; Kimi K3 offers a 1-million-token context window and is priced at USD 3 and USD 15 on OpenRouter. Thus, its 'low price' is relative to its capability tier—not necessarily the absolute lowest across all models. More importantly, open weights allow enterprises to deploy the model on their own infrastructure, regional clouds, or third-party inference platforms, reducing reliance on a single API provider. Kimi K2.5 has already accumulated over 12.3 billion tokens in usage on OpenRouter, demonstrating that Chinese open models are rapidly integrating into the global AI developer toolchain, moving beyond being merely domestic AI chatbot interfaces.
Model routing means that enterprises will no longer assign all tasks to a single model such as OpenAI or Anthropic. Instead, upon receiving a request, they will dynamically select the most suitable model based on quality, cost, latency, compliance, context length, and availability: simple classification, summarization, and customer service tasks are delegated to low-cost open models, while complex reasoning, critical code generation, and high-stakes decisions invoke expensive frontier models.

Microsoft’s Model Router now supports three routing modes—cost, quality, and balanced—and provides automatic failover. Microsoft estimates that the router can direct 60%–80% of traffic to cheaper models without measurable degradation in quality. Internal AWS tests show that intelligent routing across different model families can reduce costs by approximately 16%–56%, with savings reaching as high as 63.6% in certain RAG (Retrieval-Augmented Generation) benchmarks. This signals a shift in the core entry point for AI applications—from 'model brand' to the routing and orchestration layer, which controls traffic allocation, model pricing negotiations, and actual token consumption.
This is precisely where Jevons Paradox may reemerge in the AI industry: model routing, caching, quantization, and open weights lower the marginal cost per inference, but inexpensive intelligence will catalyze new demand previously uneconomical—such as perpetually running enterprise agents, hundreds of internal reasoning steps with self-verification, parallel sub-agents, million-token document analysis, real-time audio and video understanding, and background AI embedded into every software operation. As long as token demand exhibits price elasticity greater than one—meaning that a 50% drop in unit price leads to more than a doubling of usage—total compute consumption and overall inference spending will actually increase.
Thus, while low-cost models may temporarily fuel concerns about 'peak compute demand,' they are more likely in the long run to diffuse AI from a handful of high-value tasks into billions of routine workflows, transforming the capital expenditures of the training era into sustained, distributed, and high-frequency compute consumption in the inference era. The open-weight joint letter itself defines economic sustainability as reserving frontier capabilities exclusively for frontier problems, while delegating vast volumes of everyday tasks to efficient, specialized models.
AI chips, storage, optical interconnects, AI control planes, and vertical applications
For the AI compute supply chain, this does not imply a disappearance of AI compute demand, but rather a structural shift from centralized training toward a balanced emphasis on both training and widespread inference. Frontier training still requires the most advanced GPUs, leading-edge semiconductor processes, HBM/enterprise-grade DRAM/data center NAND, advanced packaging, and high-speed optical interconnects. Meanwhile, open models and model routing will expand inference deployments across enterprise private clouds, sovereign AI infrastructures, regional clouds, and local data centers, driving demand for inference-optimized GPUs, specialized accelerators, server CPUs, high-end DRAM/NAND memory components, Ethernet infrastructure, optical interconnect systems such as transceivers, and power and liquid cooling solutions.
Moreover, the technological pricing anchor is shifting from a sole focus on peak FLOPS toward metrics such as tokens per watt, tokens per dollar, memory bandwidth, KV cache efficiency, MoE sparse activation, and cluster utilization. It must be emphasized that open weights do not equate to low-cost operation on personal computers: Kimi K3, with its 2.8 trillion total parameters and 1 million-token context window, still requires deployment at the scale of large server racks or data centers—a scenario that favors vendors with capabilities in large-scale inference, cluster scheduling, and energy infrastructure.
OpenAI and Anthropic will not lose all growth potential, but the 'token tollbooth' model—where every request invokes expensive closed-source models—will face structural erosion. Anthropic’s Claude Opus 5 is priced at $5 and $25 per million input and output tokens, respectively, while Sonnet 5 is offered at promotional rates of $2 and $10. Faced with open models like Kimi and automated routing, closed-source labs must demonstrate that their frontier capabilities, reliability, security compliance, tooling ecosystems, and end-to-end agent performance deliver business value significantly exceeding the price differential.
In Goldman Sachs’ view, the AI computing super bull market is far from over; it is transitioning from the initial phase of an ‘AI chip buying frenzy’ into a second stage characterized by ‘large-scale construction of AI factories.’ Consequently, the next wave of excess alpha returns will no longer be confined solely to leading players in AI GPUs and AI ASICs but will systematically spread across the entire stack of AI computing infrastructure—including high-performance data center CPUs, DRAM/NAND/HBM memory, AI PCBs, liquid cooling systems, data center optical interconnects, ABF substrates/glass core substrates, MLCCs, electronic fabrics, and broad-based wafer foundry services.
A recent research report led by Brian Nowak, senior analyst at Morgan Stanley, one of Wall Street’s financial titans, has again significantly raised its capital expenditure forecasts for the world’s five largest hyperscale cloud providers and manufacturers (Meta, Amazon, Microsoft, Google, and SpaceX), projecting approximately $1.2 trillion for 2027 and $1.4 trillion for 2028. The firm also revised its 2026 capital expenditure outlook for major U.S. tech giants upward dramatically—from $433 billion a year ago to $805 billion.

In this latest update, Morgan Stanley raised its 2027 and 2028 capital expenditure forecasts for Meta by 29% and 22%, respectively, to $225 billion and $250 billion; for Amazon, it increased the corresponding forecasts by 15% and 29%, to $308 billion and $318 billion. Morgan Stanley stated that the capex supercycle is not yet over, with 2026 and 2027 likely representing the steepest years of growth. Beyond 2028, stock performance will hinge less on ‘who spends the most’ and more on ‘who can most rapidly convert AI computing resources into revenue, profit, and free cash flow.’
Although MoE, model routing, quantization, and caching reduce the cost per token, they may—through Jevons Paradox—stimulate more persistent agents, longer inference chains, and higher invocation frequencies, thereby causing total compute and total power consumption to continue rising.
A latest research report from Citigroup, another Wall Street financial giant, indicates that the AI-era arms race is shifting from 'whose model is the smartest' to 'who can sustainably produce intelligence at the lowest cost and highest efficiency under physical constraints.' Open-weight models like Kimi K3 are rapidly approaching the frontier of closed-source models, accelerating the commoditization of model capabilities. However, simultaneous growth in parameter count, context length, and multi-step agent reasoning is shifting bottlenecks away from raw FLOPs toward HBM capacity and bandwidth, GPU interconnect speeds, cluster scheduling efficiency, and power availability. NVIDIA research also notes that as model size, sequence length, and batch sizes grow, HBM often becomes the primary scaling constraint. Meanwhile, the International Energy Agency (IEA) forecasts that electricity demand from AI data centers is growing significantly faster than overall power demand, while grid infrastructure typically takes longer to build than data centers themselves.
If the AI super bull market continues, investment focus in the AI computing power supply chain may remain concentrated over the long term on AI chips, memory/storage, optical interconnects, the AI control plane, and vertical applications. This explains why Wall Street financial giants such as Morgan Stanley, Bank of America, and Nomura remain consistently bullish on leading players in the AI computing power ecosystem—including NVIDIA, Micron, SK Hynix, and Intel—namely, those who dominate foundational AI infrastructure such as AI chips, memory components, networking infrastructure, and energy-constrained compute resources; those commanding the AI control plane encompassing routing, evaluation, inference optimization, security, and observability; and those vertical applications capable of leveraging token deflation to scale usage and generate real cash flows. The greatest caution should be exercised toward software companies lacking proprietary data, workflow moats, or distribution channels, and which merely wrap a single closed-source API—because as model routing matures, their products become increasingly substitutable, and their profit margins are ever more vulnerable to being eroded by price wars at the infrastructure layer.
Editor/Deng