SemiAnalysis: The migration of wealth along the AI value chain, from infrastructure to the model layer, is accelerating.
```
The value center of the AI industry is undergoing a structural shift.
Over the past two years, Nvidia, memory manufacturers, and energy suppliers have dominated the distribution of AI investment returns. However, as Agentic AI commercialization accelerates, profit margins in the model layer are expanding at an unprecedented rate, while Nvidia and TSMC, which control the power supply side, have yet to fully reflect this trend in their pricing.
Anthropic is the most direct footnote to this transformation. According to the latest research by SemiAnalysis, Anthropic's annualized revenue (ARR) has soared from $9 billion at the beginning of the year to over $44 billion, with the gross margin of its inference infrastructure surging from 38% to over 70% in the same period. Meanwhile, the cost of token production has been drastically reduced due to hardware iteration and software optimization, further widening the gap between value and cost and pushing model providers into a new phase of rapidly rising profit margins.
On the supply side, Nvidia and TSMC possess the most scarce resources, but have yet to respond adequately to the current surge in demand through pricing. SemiAnalysis believes this lag in pricing constitutes a significant market dislocation: next-generation systems represented by Vera Rubin (VR NVL72) have substantial room for price increases. Whoever seizes the initiative in this wave of value redistribution will profoundly influence the investment logic at every link of the AI industry chain.
Three-Year Migration Path of the AI Value Pool
From 2023 to 2025, the excess returns from AI investment are mainly concentrated at the infrastructure layer.
In May 2023, Nvidia released a blockbuster earnings report, soaring 25% in after-hours trading in a single day and officially igniting the AI investment wave. In 2024, Vistra and GE Vernova rose 265% and 146% respectively, becoming the top-performing S&P 500 stocks as the energy bottleneck became a market focus. In 2025, the memory sector took the lead, with SanDisk, Western Digital, Seagate, and Micron all recording annual increases of over 200%, and the imbalance between storage supply and demand became the core variable driving pricing.
Meanwhile, model providers and inference service vendors faced long-term pressure on gross margins. At the time, skeptics argued that AI's utility was merely a "better Google search" with a chat interface, which was a far cry from the anticipated multi-trillion-dollar capital expenditures.
This pattern changed fundamentally at the end of 2025.
Agentic AI: The Inflection Point of Token Economics Reshaping
SemiAnalysis sees December 2025 as the true inflection point for AI commercialization—Agentic AI begins to operate stably and is widely deployed in enterprise workflows. The core significance of this change is that it fundamentally alters the economic value of tokens.
Take SemiAnalysis itself as an example, its annualized token spending is already equivalent to about 30% of its total employee compensation, and each employee consumes over 5 billion tokens per month, more than five times the per capita level within Meta. The research team cited many real cases: Financial modeling, chart creation, and profitability analysis that previously required hours of work by junior analysts can now be completed by agents at extremely low token cost, whereas equivalent labor costs used to be hundreds to thousands of dollars.
Meanwhile, the cost of token production is plummeting. SemiAnalysis estimates that in agent task scenarios, the actual blended price for running Opus 4.7 is about $0.99 per million tokens, far lower than the official prices of $5/$25—because agent workloads have a very high input-output ratio (about 300:1) and a cache hit rate of over 90%, so a large number of tokens fall into the lowest price tier.
The acceleration at the hardware level is also significant. Compared with last year's H100, the Blackwell series can generate about 30 times more tokens per second under frontier workloads. Further comparison shows that, in the most optimized configuration, the GB300 NVL72 has about 17 times higher throughput under FP8 precision than the most optimized H100, and this gap increases to 32 times under FP4, while the total cost of ownership (TCO) is only about 70% higher.

The twin scissors difference between value and cost is precisely the driving force behind Anthropic's gross margin leap from 38% to over 70%.
Pricing Power in the Model Layer: Why It Won't Be Eroded By Competition
In the face of the rapid expansion of model providers' profit margins, the most common market doubt is: Competition will eventually drive down prices. SemiAnalysis takes a reserved stance and offers two supporting points.
First, the pricing power of leading closed-source models remains solid. Although open-source models continue to improve in benchmark tests, their performance in real knowledge work scenarios is still significantly weaker than that of leading closed-source models. For example, Kimi K2.6 (priced at $0.95/$4) exerts very limited downward pressure on Anthropic Opus's pricing.
Second, compute constraints mean that no single frontier lab can meet the market's total demand alone. Anthropic has started to actively manage demand by locking Claude Code behind a $100+ monthly subscription and limiting third-party access. Token demand is expected to continue exceeding supply for the foreseeable future. This structural scarcity gives leading model providers the confidence to price based on value rather than cost.
Anthropic has implemented this logic through its product line strategy: The Opus fast SKU is priced 6x higher than standard Opus, and the soon-to-launch Mythos is priced at $25/$125, 5x the price of regular Opus, and leading enterprise customers are still willing to pay for these high-priced SKUs. SemiAnalysis says that if Anthropic prices Mythos fast at $150/$750, they themselves would also be a paying user.
Nvidia & TSMC: Pricing Lag of Scarce Resources
However, the two companies holding the most core scarce resources—Nvidia and TSMC—have yet to fully catch up with this wave of value reappraisal.
TSMC's N3 advanced process capacity has become the tightest bottleneck for the expansion of AI computing power. Nvidia, Broadcom, Annapurna, MediaTek, and AMD are all competing for limited N3 wafer quotas, while N3 capacity utilization is expected to exceed 100% in the second half of 2026. DRAM fab utilization is over 90%, and overall memory supply remains tight, but pricing remains relatively conservative.
SemiAnalysis believes that TSMC is fully qualified to raise prices substantially, and customers would not only accept this, some would even welcome it—Nvidia is a typical example: If TSMC raises prices, competitors would get fewer wafer quotas, and Nvidia paying higher wafer prices would actually help strengthen its market position. Nvidia CEO Jensen Huang publicly stated in 2024 that TSMC should increase wafer prices, with this logic in mind.

Nvidia's own pricing strategy is similarly conservative. SemiAnalysis points out that Nvidia's pricing framework is still anchored on the previous assumption that "the unit price users are willing to pay for compute power declines over time," but this assumption is no longer valid. As agent workloads take off, demand for compute is no longer growing linearly, but instead is showing compound acceleration.
Rubin System: Quantifying Nvidia’s Pricing Headroom
Using the Vera Rubin (VR NVL72), which will be released in the second half of 2026, as a reference, SemiAnalysis has built a "One Chart to Rule Them All" pricing analysis framework, anchoring the floor and ceiling for rental pricing from cost and value perspectives, respectively.

Cost side (floor): Based on the deployment threshold of new cloud (Neocloud) providers requiring an internal rate of return (IRR) of no less than 15.6%, the minimum rental price for VR NVL72 should be about $4.92 per GPU per hour to maintain deployment willingness.
Value side (ceiling): Using the current 5-year contract rental of the GB300 at $0.70 per PFLOP as a benchmark, the corresponding rental cap of VR NVL72 is about $12.25 per GPU per hour.

Currently, the pricing of the VR NVL72 system puts the per PFLOP cost at only about $0.28, marking a 60% decrease compared to the GB300 NVL72 and far exceeding the historic improvement trend. This means Nvidia has about 40% price increase headroom for its servers— even after a price hike, Neoclouds would still retain enough profit, and the overall cost improvement rate remains below historic trends.
SOCAMM memory pricing is another key variable. The VR NVL72 uses slot-in LPDDR5X memory modules (SOCAMM), which can be priced independently of compute units. SemiAnalysis estimates that Nvidia will pay about $8 per GB for SOCAMM contracts in Q1 2026, a significant jump from the previous quarter; and by the end of 2026, SOCAMM prices may exceed $13 per GB. Against this backdrop, it would be rational for Nvidia to target a 60% gross margin on SOCAMM: on one hand, memory supply is constrained while Nvidia holds the largest share; on the other hand, the VR NVL72's leading TCO performance leaves customers without alternatives.
Value Destination: Who Is Winning and Who Is Waiting
SemiAnalysis's framework reveals the core contradiction in current AI value distribution: Improvements in token economics are rapidly boosting profits for model providers, inference service providers, and Neocloud, but on the supply side—holders of the most scarce compute resources, Nvidia and TSMC—there’s a clear misalignment between their pricing behavior and their supply-side scarcity.
This misalignment is essentially a proactive choice—Nvidia is acting like an "AI central bank", passing value downstream by improving software efficiency to maintain the ecosystem's long-term expansion drive, while also avoiding antitrust regulatory pressure. TSMC continues its longstanding pricing philosophy of "stabilizing the ecosystem rather than extracting all the upside".
However, as inference ROI becomes increasingly clear and value-based pricing logic becomes mainstream, the pressure for these two companies to switch to a value-based pricing framework will continue to mount. Once that happens, the value distribution pattern of the AI supply chain will be reshaped again—at that point, bargaining power on the compute supply side will return more strongly to the hardware layer.
Risk Warning and DisclaimerThe market involves risk; investment must be cautious. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their particular circumstances. Invest accordingly at your own risk. ```