share_log

The AI price war is intensifying while barriers to entry are declining: How long can substantial capital expenditures be sustained?

wallstreetcn ·  Sep 16 16:55

Jefferies believes that advances in model‑training techniques have lowered the barrier to entry for large‑scale models, enabling new players to break into the top tier with funding rounds in the tens of millions of dollars, while intensifying competition on API pricing. Although demand for AI inference continues to grow and computing power needs remain on the rise, narrowing performance gaps and declining costs are increasingly putting the ability of substantial capital expenditures to deliver adequate returns to the test.

A new dynamic is emerging in the AI large‑model competition: model supply continues to expand, and techniques such as model distillation are steadily lowering the R&D barrier, enabling new entrants to tap into the high‑performance model market with capital investments far below those of the leading players. According to Jefferies, this trend is driving down API pricing and challenging the industry's longstanding reliance on massive capital expenditures to erect competitive barriers.

According to a research report released by Jefferies on September 16, the global top-14 list of language models tracked by Artificial Analysis (AA) added four new models in September, originating from Singapore, South Korea, the United Arab Emirates, and the United States. The research institutions behind these models generally have limited funding; some have raised only about $20 million in total yet have already made it onto the global leaderboard of leading models.

Jefferies analyst Edison Lee and colleagues argue that the growing number of players in the LLM market, coupled with increasing regional support for homegrown models and the widespread adoption of model‑distillation techniques, is further lowering barriers to entry. At the same time, narrowing performance gaps among models and intensifying price competition for APIs are imposing new pressures on the AI supply chain—characterized by "high capital expenditures and low‑certainty returns."

New players can enter the market at low cost, and distillation technology lowers the barrier to entry.

The report indicates that in September, the four models newly ranked within the top 14 of the AA list are Agnes 3.0 Flash from Singapore's Sapiens AI, Motif-3-Beta from South Korea's Motif Technologies, K2 Horizon from the UAE's MBZUAI, and Apodex 1.1, which is headquartered in California, USA, with a research team based in Singapore.

Among them, Sapiens AI has raised approximately US$20 million in cumulative funding, while Motif Technologies has closed a Series B round of about US$16.6 million; overall investment levels still lag significantly behind those of leading AI companies.

Such cases demonstrate that the capital threshold for large‑model competition is declining. Some new entrants are bypassing the traditional approach of building and training massive foundational models from scratch; instead, they leverage open‑source models, smaller teams, and more efficient training techniques to iterate rapidly.

Distillation techniques are a key driver in this space. By leveraging the outputs of existing large models to train new ones, developers can acquire some of the capabilities of advanced models at a lower cost. Jefferies believes that such methods are difficult to fully block, and as their adoption continues to grow, the cost for latecomers to catch up with leading models may continue to decline.

However, low‑cost entry does not mean that the leading players have lost their advantage. Training state‑of‑the‑art models still requires substantial computational resources, data, and engineering effort, with new entrants typically gaining competitiveness only in specific capabilities or niche use cases. The real shift lies in the widening capital barrier between "entering the market" and "becoming the strongest."

API pricing is under pressure, and models are beginning to enter price competition.

The increase in the number of models has begun to feed through to API pricing. According to Jefferies data, the API prices for some new models are already significantly lower than those of leading models; for instance, Apodex 1.1's hybrid API price stands at $0.30 per million tokens, making it one of the most competitively priced among the top 15 models.

Meanwhile, leading vendors have not universally cut prices; instead, they are pursuing higher pricing by enhancing the capabilities of their next-generation models. The report argues that such price hikes not only reflect improvements in model performance but may also serve as a way to demonstrate commercialization prowess to capital markets—leveraging greater intelligence to justify higher price points and thereby bolster market expectations for profit margins and investment returns.

The issue is that price hikes and price cuts are happening simultaneously. Leading models are attempting to maintain their premium pricing through performance upgrades, while new entrants are vying for developers and application demand by offering lower‑cost solutions. According to Jefferies, the U.S. market alone currently hosts seven major large‑model providers; this growing number of players suggests that price competition in the API market could persist.

From "competing on intelligence" to "competing on efficiency"

Another shift in model competition is that the industry is moving from solely pursuing model intelligence to simultaneously benchmarking inference efficiency. This month, Jefferies introduced the AutomationBench-AA benchmark to assess models' performance on agent-based automation tasks and has accordingly recalibrated historical data.

Against this backdrop, some model providers have begun reducing inference costs through optimizations to model architectures and memory usage. On September 10, DeepSeek released its Flash model version 4.1, lowering the hybrid API pricing to $0.20 per million tokens—a 74% reduction compared to the previous version.

The model further reduces hardware requirements through techniques such as a causal encoder–decoder architecture, an engram‑based conditional memory mechanism, and Compressed Sparse Attention 2: HBM demand is cut by roughly one quarter, while SSD requirements are reduced by about one eighth. Jefferies believes that this approach places greater emphasis on computational and memory efficiency, rather than simply maximizing model intelligence.

This also means that the competitive landscape for AI models is expanding. As demand for inference tokens grows rapidly, model providers must not only enhance model capabilities but also reduce the cost of generating each token. For application developers, if lower‑cost models can already handle a sufficient range of tasks, price differences between models may become more significant than performance disparities alone.

Capital expenditures are still rising, but returns need to be revalidated.

Jefferies' concerns about the AI supply chain center not on a decline in AI demand, but on whether the pace of capital investment can keep up with expected returns. On the one hand, new entrants are now able to secure funding in the tens of millions of dollars to break into the ranks of the world's leading large‑scale models; on the other, incumbent leaders continue to ramp up investments in computing power, data centers, and infrastructure.

From the demand side, AI inference demand continues to grow rapidly, and improvements in model capabilities are constantly giving rise to new application scenarios, meaning that the underlying drivers of computing power demand remain intact. The key question is: as model offerings become increasingly diverse and API pricing comes under competitive pressure, how long will it take for additional infrastructure investments to translate into sufficient revenue and profitability?

Jefferies therefore believes that the AI industry may gradually enter a phase that places greater emphasis on capital efficiency. For leading companies, continuing to ramp up capital expenditures may still be a key strategy for maintaining technological leadership; however, the industry as a whole will need to pay closer attention to redundant infrastructure investments, resource utilization, and the tangible commercial value of newly added computing capacity.

Ultimately, Jefferies expects the industry to reduce overall capital expenditures through consolidation, partnerships, and more efficient resource allocation. While demand for AI infrastructure remains robust, "demand growth" does not automatically translate into "sufficient returns on all capital investments," a factor that will be critical as the market reassesses AI valuations in the next phase.

Editor/Deng

The translation is provided by third-party software.


The above content is for informational or educational purposes only and does not constitute any investment advice related to EleBank. Although we strive to ensure the truthfulness, accuracy, and originality of all such content, we cannot guarantee it.