share_log

SemiAnalysis: Wealth is shifting at an accelerating pace across the AI value chain, from infrastructure to the model layer.

wallstreetcn ·  Jun 29 11:06

The commercialization of AI agents is accelerating, driving surging profits for model developers. Meanwhile, NVIDIA and Taiwan Semiconductor—the companies holding the scarcest computing power—suffer from significant 'pricing lag.' This market misalignment implies substantial room for price increases, setting the stage for a major revaluation of the AI industry chain and an impending reshuffling of wealth.

The value center of the AI industry is undergoing a structural shift.

Over the past two years, NVIDIA, memory manufacturers, and energy suppliers have dominated the distribution of AI investment returns, but as the commercialization of Agentic AI accelerates, profit margins at the model layer are expanding at an unprecedented pace, while those controlling the compute supply side $NVIDIA (NVDA.US)$ and $Taiwan Semiconductor (TSM.US)$ have yet to fully reflect this trend in their pricing.

Anthropic is the most immediate illustration of this shift. According to the latest research from SemiAnalysis, Anthropic’s annualized recurring revenue (ARR) has surged from $9 billion at the beginning of the year to over $44 billion, while the gross margin of its inference infrastructure has jumped from 38% to over 70% during the same period. Meanwhile, token production costs have been significantly compressed due to hardware iterations and software optimizations, causing the scissors gap between value and cost to widen continuously and propelling model providers into a new phase of rapidly rising profitability.

On the supply side, NVIDIA and Taiwan Semiconductor command the scarcest resources but have yet to respond adequately with price adjustments to the current surge in demand. SemiAnalysis argues that this pricing lag represents a significant market misalignment: next-generation systems such as Vera Rubin (VR NVL72) possess substantial room for price increases, and whichever players seize the initiative in this reallocation of value will profoundly influence investment logic across all segments of the AI value chain.

The Three-Year Migration Path of AI Value Pools

Between 2023 and 2025, excess returns from AI investments were primarily concentrated in the infrastructure layer.

$NVIDIA (NVDA.US)$ It first released a blockbuster earnings report in May 2023, triggering a 25% after-hours surge in a single day and officially launching the AI investment wave. In 2024, Vistra and GE Vernova rose by 265% and 146%, respectively, becoming the top-performing stocks in the S&P 500, with energy bottlenecks emerging as a key market focus. In 2025, the memory sector took the lead, $SanDisk (SNDK.US)$$Western Digital (WDC.US)$$Seagate Technology (STX.US)$and$Micron Technology (MU.US)$ all recorded annual gains exceeding 200%, with storage supply-demand imbalances becoming the core driver of pricing dynamics.

Meanwhile, model developers and inference service providers faced prolonged pressure on gross margins. At the time, critics dismissed AI’s practical utility as merely “a better Google search” wrapped in a chat interface—a perception starkly at odds with the trillions of dollars in anticipated capital expenditures.

This dynamic underwent a fundamental transformation by the end of 2025.

Agentic AI: The Inflection Point Reshaping Token Economics

SemiAnalysis identifies December 2025 as the true inflection point for AI commercialization—when agentic AI begins operating reliably and achieves large-scale deployment within enterprise workflows. The core significance of this shift lies in its fundamental transformation of the economic value of tokens.

Taking SemiAnalysis itself as an example, its annualized token expenditure now amounts to approximately 30% of total employee compensation costs, with each employee consuming over 5 billion tokens per month—more than five times Meta’s internal per-employee average. The research team cites multiple real-world cases: tasks such as financial modeling, chart generation, and earnings analysis, which previously required junior analysts several hours to complete, can now be executed by AI agents at minimal token cost, whereas the equivalent human labor would have cost hundreds to thousands of dollars.

Simultaneously, token production costs are plummeting. SemiAnalysis estimates that in agentic task scenarios, the effective blended price for running Opus 4.7 is approximately $0.99 per million tokens—far below the official list prices of $5 or $25. This is because agentic workloads exhibit an extremely high input-to-output ratio (around 300:1) and a cache hit rate exceeding 90%, resulting in the majority of tokens falling into the lowest pricing tier.

Hardware-level acceleration is equally pronounced. Compared to the H100 from a year ago, the Blackwell series delivers roughly a 30-fold increase in tokens generated per second under cutting-edge workloads. Further comparisons show that an optimally configured GB300 NVL72 achieves approximately a 17-fold throughput improvement over an optimized H100 at FP8 precision; this gap widens to 32-fold when switching to FP4 precision, while total cost of ownership (TCO) increases by only about 70%.

This dual divergence between value and cost is the primary driver behind Anthropic’s gross margin surge from 38% to over 70%.

Pricing Power at the Model Layer: Why It Won’t Be Eroded by Competition

In response to the rapid expansion of model vendors’ profit margins, the most common market skepticism is that competition will inevitably drive prices down. SemiAnalysis remains skeptical of this view and offers two supporting arguments.

First, pricing power for leading closed-source models remains solid. Although open-source models continue to set new benchmarks in standardized evaluations, their real-world performance in knowledge-intensive tasks still lags significantly behind that of state-of-the-art closed-source models. For instance, Kimi K2.6 (priced at $0.95/$4) exerts only limited downward pressure on Anthropic’s Opus pricing.

Second, compute constraints mean no single frontier lab can independently meet the entire market’s demand. Anthropic has already begun actively managing demand—for instance, by locking Claude Code behind subscription tiers priced above $100 per month and restricting third-party access. Token demand is expected to persistently outstrip supply in the foreseeable future. This structural scarcity empowers leading model providers to price based on value rather than cost.

Anthropic has already operationalized this logic through its product lineup: Opus Fast is priced at six times the standard Opus rate, and the upcoming Mythos is priced at $25/$125—five times the standard Opus rate—with top-tier enterprise clients still willing to pay for these premium SKUs. SemiAnalysis notes that if Anthropic were to price Mythos Fast at $150/$750, it would become a paying customer itself.

NVIDIA and Taiwan Semiconductor: Pricing Lags Behind Scarce Resource Value

However, the two companies that control the most critical scarce resources—NVIDIA and Taiwan Semiconductor—have not fully kept pace with this wave of value reassessment.

Taiwan Semiconductor’s N3 advanced-node capacity has become the tightest bottleneck in the entire AI compute expansion. NVIDIA, Broadcom, Annapurna, MediaTek, and AMD are all competing for limited N3 wafer allocations, and N3 capacity utilization is expected to exceed 100% in the second half of 2026. DRAM wafer fab utilization has already surpassed 90%, indicating overall tight memory supply, yet pricing remains relatively conservative.

SemiAnalysis believes Taiwan Semiconductor is well positioned to significantly raise prices, and customers would not only accept it but some might even welcome it. NVIDIA serves as a prime example: if a TSMC price hike means competitors receive smaller capacity allocations, NVIDIA paying higher wafer prices could actually reinforce its market leadership. In 2024, NVIDIA CEO Jensen Huang publicly stated that TSMC should increase wafer prices—a view grounded precisely in this strategic logic.

NVIDIA’s own pricing strategy exhibits a similarly conservative tendency. SemiAnalysis notes that NVIDIA’s pricing framework remains anchored to the prior assumption that 'willingness to pay per unit of compute declines over time,' an assumption that no longer holds. With the explosion of agent-driven workloads, compute demand is no longer growing linearly but accelerating at a compound rate.

The Rubin System: Quantifying NVIDIA’s Pricing Headroom

Using the Vera Rubin (VR NVL72), scheduled for release in the second half of 2026, as a benchmark, SemiAnalysis has developed a 'One Chart to Rule Them All' pricing analysis framework that anchors rental pricing between a floor derived from cost considerations and a ceiling based on value creation.

Cost Floor: Based on the deployment threshold that Neocloud (emerging cloud service providers) must achieve an internal rate of return (IRR) of at least 15.6%, the minimum hourly rental rate per GPU for the VR NVL72 must be approximately $4.92 to sustain Neocloud’s willingness to deploy.

Value Ceiling: Anchored to the current five-year contract rental rate for the GB300 of approximately $0.70 per PFLOP, the implied upper bound for VR NVL72 rental pricing is about $12.25 per GPU per hour.

Currently, the VR NVL72 system pricing reduces the cost per PFLOP to approximately $0.28, representing a 60% decline compared to the GB300 NVL72—far exceeding historical trendline improvements. This implies NVIDIA has roughly 40% room to raise server prices; even after such an increase, sufficient profit margins would remain for Neocloud, and the overall cost improvement would still fall below historical norms.

SOCAMM memory pricing represents another critical variable. The VR NVL72 employs socketed LPDDR5X memory modules (SOCAMM), which can be priced independently from the compute units. SemiAnalysis estimates that NVIDIA’s contracted SOCAMM price in Q1 2026 was approximately $8 per GB, marking a significant jump from the previous quarter; by the end of 2026, SOCAMM prices could exceed $13 per GB. Against this backdrop, NVIDIA achieving a 60% gross margin on SOCAMM appears logically reasonable: on one hand, memory supply is constrained and NVIDIA commands the largest market share; on the other, the VR NVL72’s performance leadership at the total cost of ownership (TCO) level leaves customers with few viable alternatives.

Value Capture: Who Is Winning, and Who Is Waiting

SemiAnalysis’s framework reveals a core contradiction in the current AI value distribution: improvements in token economics are rapidly boosting profits for model providers, inference service vendors, and Neoclouds, yet there is a clear misalignment between the pricing behavior of NVIDIA and Taiwan Semiconductor—the controllers of the most scarce resource on the compute supply side—and the scarcity of their supply.

This misalignment is, in essence, an active strategic choice. NVIDIA is assuming a role akin to an “AI central bank,” channeling value downstream through software efficiency gains to sustain long-term ecosystem expansion while mitigating antitrust regulatory pressure. Meanwhile, Taiwan Semiconductor continues its historical pricing philosophy of stabilizing the ecosystem and refraining from capturing the full upside of market booms.

However, as return on investment (ROI) for inference becomes increasingly clear and value-based pricing logic gains broader market acceptance, both companies will face mounting pressure to transition toward a value-based pricing framework. Once this shift occurs, the AI industry’s value distribution landscape will be reshaped once again—with bargaining power on the compute supply side shifting more decisively back toward the hardware layer.

Editor/KOKO

The translation is provided by third-party software.


The above content is for informational or educational purposes only and does not constitute any investment advice related to EleBank. Although we strive to ensure the truthfulness, accuracy, and originality of all such content, we cannot guarantee it.