Rubin shrinks, the CUDA myth of Nvidia is starting to falter.
```
Two seemingly unrelated pieces of news landed one after another in the last week of June.
On June 25, OpenAI released its first self-developed AI inference chip, Jalapeño, partnering with Broadcom to complete the process from design to tape-out in just nine months—the world's largest GPU buyer is now making its own chips.
On June 30, semiconductor research institute SemiAnalysis publicly announced on social platforms: NVIDIA’s original 4-chip Rubin Ultra was canceled only three months after its GTC 2026 launch, with the new version’s performance cut by nearly half. "The context for all this," the institute added, "is that NVIDIA’s market share is being eroded."
As early as last year, media reported that Anthropic’s annualized revenue was approaching $7 billion, with Claude Code generating $500 million in annualized income within two months of launch. The compute base driving all this is no longer just NVIDIA—Google TPU handles training, Amazon Trainium does inference, and NVIDIA GPU has become the third option for ‘research and exploration’.
Three pieces of news, all pointing to the same issue: CUDA’s moat—NVIDIA’s most solid yet most mythologized competitive barrier—is beginning to crack.
From 87% to 75%, NVIDIA's "irreplaceable" status is collapsing
Let's look at some numbers.
According to Silicon Analysts' estimates based on NVIDIA/AMD earnings and TSMC capacity data, NVIDIA’s share in the AI accelerator market (by revenue) follows this trajectory:

We can see that NVIDIA’s revenue continues to grow—from $15 billion to $150 billion, tenfold in four years. But its market share has slid from the peak of 87% to 75%, meaning a growing portion of the incremental market is being carved away.
That lost portion isn’t taken by a single rival, but by competition from all sides: Google TPU, Amazon Trainium, Microsoft Maia, Meta MTIA, Broadcom’s custom XPU—and now, OpenAI.
Broadcom CEO Hock Tan revealed a previously undisclosed number on the FY2026 Q1 earnings call: Broadcom’s AI semiconductor annualized run rate has reached $8.4 billion, up 106% year-on-year and aiming for a $40-50 billion annual trajectory. The company has signed six hyperscale customers for custom AI chips, OpenAI being the sixth.
In other words, the world’s largest cloud and AI companies have unanimously chosen the same direction: making their own chips.
Anthropic’s Choice
If market share data is cold statistics, Anthropic's case is a living textbook of “de-NVIDIAization”.
Anthropic is one of the world's fastest-growing AI companies. Annualized revenue is approaching $7 billion (just about $1 billion the same period in 2025), serving over 300,000 enterprise customers, with large client numbers up 7x year-on-year. Claude Code achieved $500 million annualized revenue within two months of launch, termed by Anthropic as "the fastest growing product in history."
The compute base driving all this is a three-platform architecture that Anthropic CFO Krishna Rao calls a "unique compute strategy":

Note the last column. NVIDIA GPU ranks third—not tied, not an alternative, but the smallest of the three options.
This is not a cash-strapped small company improvising with cheap substitutes. It’s the world’s second-largest AI company, using non-NVIDIA chips in production to drive its fastest-growing products.
SemiAnalysis pointed this out specifically in its June 30 post: “A considerable portion of Claude Code’s inference runs on Trainium; Claude’s training is done on TPU. Just a year ago, TPU and Trainium reaching this scale, with CUDA’s moat slowly being eroded, was unimaginable.”
Why is Anthropic doing this? Not because TPU or Trainium are stronger than H100—in absolute performance, they may still lag. But in specific scenarios, proprietary chips are far more cost-effective than general-purpose GPUs. TPU for training, because Google gave a multi-billion-dollar contract and promised supply of millions of chips. Trainium for inference, because AWS is its main cloud provider, having invested $8 billion, and Project Rainier’s supercomputing clusters run entirely on Trainium 2, with no GPU premium.
Amazon is betting big on Trainium. Its FY2026 Q1 disclosure showed the Trainium line has secured over $225 billion in revenue commitments, including from OpenAI and Anthropic. AWS AI revenue run rate now exceeds $15 billion, with most Bedrock inference running on Trainium.
The key here is not “performance,” but “cost.” Inference burns money every day. Each ChatGPT reply, every API code return, means a GPU running on power behind the scenes. Anthropic’s use of Trainium instead of GPU for inference isn’t to be faster, but to get more compute per dollar spent.
Three points of erosion: Where CUDA's moat cracks
CUDA is seen as NVIDIA’s strongest moat because it built a closed ecosystem of "hardware-software-developer":
- 20 years of accumulation, 4+ million developers
- All mainstream ML frameworks optimized for CUDA as priority
- cuDNN, TensorRT, NCCL, etc. form deep binding
- Switching costs measured in years and hundreds of millions of dollars
But in 2026, AI chip competition isn’t about “making a GPU 10% faster than H100”—that’s a frontal assault no one can win. Erosion comes from three sides:
Erosion path one: Self-developed ASICs—not fighting on all battlefields, just targeting the fattest inference cake
This is the deadliest path. The logic is not “I can do better than NVIDIA,” but “I don’t need all the features of a GPU, just inference.”
A single NVIDIA H100 does: graphics rendering, scientific compute, AI training, AI inference, video encoding and decoding … Jalapeño does just one thing: runs OpenAI’s models for inference. The former is a Swiss Army knife, the latter an axe specialized for one type of wood—in certain tasks, the axe is way more useful and much cheaper.
OpenAI Jalapeño’s positioning is extremely precise: don’t compete with NVIDIA on versatility, just on inference—the use case consuming billions of API calls daily, burning hundreds of millions in annual costs—to the extreme. OpenAI’s official aim is to cut inference costs by 30-50%. At a scale burning millions in inference fees daily, this means saving hundreds of millions in pure profit annually.
And OpenAI is not the first. Microsoft Maia 200 (released Jan 2026), Google TPU Ironwood (Gen 7, first inference-focused), Amazon Trainium 3—all four major clouds have launched their own inference chips. Add Meta MTIA and Apple’s custom chips, among the world’s top seven tech firms, only one is still “just buying, not making”—and it’s on its way.
Erosion path two: AMD—from “existing” to “credible alternative”
AMD's AI GPU revenue soared from under $1 billion in 2022 to over an estimated $15 billion in 2026—over 15x growth in four years.
The key turning point is the MI400 series. Based on CDNA5 architecture, 432GB HBM4 memory, 19.6 TB/s bandwidth, expected mass production in H2 2026. S&P Global forecasts MI400 alone will bring $7.2 billion in revenue, 25% of AMD’s data center business.
More importantly, client signals. Meta has committed up to 6 GW of purchases—AMD’s largest AI chip order ever, and a clear sign: hyperscale clients are hedging with multiple suppliers.
AMD's limitations are also clear: TSMC CoWoS capacity allotment is just 11%, while NVIDIA has over 60%. Capacity ceiling means AMD cannot hit NVIDIA with sheer volume in the short term. But the "credible second supplier" position alone has already torn down the “Only NVIDIA” narrative.
Erosion path three: Software decoupling—Triton, JAX, and the “CUDA-Free” Future
This is the easiest to overlook but most dangerous in the long term.
CUDA’s lock-in relies on one simple fact: AI researchers code in PyTorch, which runs on CUDA underneath. If PyTorch's backend no longer relies on CUDA?
This is happening. The PyTorch team has verified that Triton compiler enables “CUDA-Free” inference—running Llama 3 models on H100 and A100, Triton kernels match CUDA’s token throughput. In Feb 2026, Triton launched new multi-backend support, allowing the same code to compile to different hardware—AMD GPU, Intel GPU, even various ASICs.
Google’s JAX framework goes further. Designed hardware-independent from the start—the same code runs on TPU, GPU, even CPU. Anthropic chose TPU for training largely because JAX lets them switch compute platforms without rewriting model code.
What does software-layer decoupling mean? It means a new generation of AI researchers may train cutting-edge models without writing a single line of CUDA code. When developers are no longer locked in the CUDA ecosystem, the hard logic of “must buy NVIDIA” becomes a soft option of “can buy NVIDIA.”
Rubin Ultra cancelled: The watershed of physical limits
Back to the opening news. NVIDIA's 4-chip Rubin Ultra was canceled three months after launch, with SemiAnalysis saying “execution issues at the manufacturing layer are causing more market share losses.”
The technical reason is not complex. The original Rubin Ultra plan integrated 4 compute chips + 16 HBM4E memory modules in a single package using TSMC CoWoS-L process. But per Global Semi Research, the 4-chip config caused substrate warpage—the base bent in multiple directions, making chips unable to fully contact the substrate. Signal transmission failed, chips couldn't function.
TSMC’s alternative CoPoS (panel-level packaging) won’t be ready for high volume until late 2028. NVIDIA can’t wait—so the new Rubin Ultra reverted to a 2-chip design, halving performance.
The symbolic value of this incident exceeds its practical business impact.
NVIDIA will still sell every Rubin Ultra it can make. But "from 4 chips to 2 chips" exposes a deeper issue: NVIDIA’s product iteration speed is hitting the wall of physics. Bigger chips → more complex packaging → higher defect rates → delay or shrink. This curve can’t stretch infinitely.
Meanwhile, competitors are bypassing this wall by making more specialized chips, not bigger ones.
The cracks in pricing power
NVIDIA’s truly unshakable moat is not its CUDA software ecosystem, but manufacturing. Over 60% of TSMC’s advanced CoWoS packaging capacity is in its hands. That’s a physical barrier, not a software one. Rivals may write better frameworks or design more efficient ASICs—but surpassing NVIDIA in shipments means getting past TSMC capacity first.
But here lies the problem: The manufacturing moat depends on a third-party foundry. It’s not an asset NVIDIA itself controls.
NVIDIA’s 88% gross margin—H100 costs $3,320, sells for $28,000—rests on one premise: customers cannot leave it. If this changes from “can't leave” to “best value choice,” pricing power is no longer absolute.
Anthropic proved another way: don’t chase the best chip, chase the best-fit chip. Training with TPU not GPU, because Google supplied enough chips at good prices. Inference with Trainium not GPU, because AWS is a strategic shareholder and Project Rainier bypassed generic GPU premium.
When the world’s second-largest AI company relegates GPU to the smallest of three compute platforms, “must buy NVIDIA” is no longer a hard rule.
NVIDIA is still the best. No leading AI company has fully walked away from it—Anthropic keeps part of its GPU for “advanced research exploration,” OpenAI’s Jalapeño only does inference, not training, Meta's MTIA covers only recommendation and content moderation.
But from “only NVIDIA” to “NVIDIA is the most expensive, use the cheaper first,” the gap is the loss of pricing power.
The market is already repricing this possibility. Every bearish report from SemiAnalysis this year has triggered sharp swings in related sectors: early June, SOCAMM downsizing news sent Micron down 13% in one day; June 10 CPO delay controversy forced NVIDIA execs to deny rumors; June 30 Rubin Ultra cancellation reignited debate.
Behind these swings is the market struggling to answer a question it never needed to ask before: If CUDA is not irreplaceable, what is NVIDIA worth.
Risk Warning and DisclaimerThe market has risks, invest cautiously. This article does not constitute personal investment advice, nor does it take into account individual users' specific investment objectives, financial situation, or needs. Users should consider whether any opinions, viewpoints, or conclusions in this article are suitable for their particular circumstances. Investment at your own risk. ```