OpenAI develops a chip that "surpasses Blackwell's" in 9 months; does Nvidia still have a competitive advantage?
OpenAI, in collaboration with Broadcom, has released the first real-world test data for its self-developed inference chip, Jalapeño. The chip's performance significantly surpasses existing Nvidia-level systems, sparking widespread discussion in the market about the restructuring of the AI chip landscape.
According to a recent analysis by Rich Privorotsky, head of trading at Goldman Sachs, Jalapeño is not the so-called "Nvidia killer," but its emergence signifies that the AI hardware market is shifting from high concentration to diversified differentiation. This judgment is directly related to Nvidia's long-term valuation logic— not this year's shipments, but the profit structure over the next few years.
For the market, the core impact of this event has two dimensions: First, Nvidia's moat in training chips and the CUDA ecosystem is not threatened in the short term; second, the high-frequency, high-volume daily workload of inference is becoming the main battleground for custom chips, and Nvidia's "annuity-style" revenue source is facing gradual erosion.
Jalapeño's real-world test data: Leading inference performance, but with a clearly defined positioning.
OpenAI and Broadcom recently released the first batch of publicly available benchmark test data for their jointly designed Jalapeño chip. The test results show that the chip can perform 1.5 to 1.9 times more AI computations per watt than comparable NVIDIA systems, and its response time is 1.7 to 3.6 times faster.
In terms of hardware specifications, a single Jalapeño chip has a rated power consumption of approximately 700 watts, and its actual operating power consumption is usually even lower. A complete rack can accommodate 128 chips, equipped with approximately 27.5TB of HBM4 high-bandwidth memory, with each chip having a memory bandwidth of 15.4TB/s.
It is worth noting that OpenAI did not compromise on memory configuration—this is a chip that deliberately pursues a bandwidth-intensive design, rather than cutting costs by reducing memory.
Jalapeño's deployment plan is to launch on a small scale later this year, with mass production ramping up in 2027. OpenAI has made it clear that the chip is a complement to, rather than a replacement for, its partners such as Nvidia.
The fundamental difference between training and reasoning: Where is Nvidia's competitive advantage?
The key to understanding this incident lies in distinguishing between the two core workloads of AI chips.
Training is like "going to school"—feeding massive amounts of data to the model and updating weights. This task is complex, large-scale, and time-consuming, and still heavily relies on general-purpose GPUs like NVIDIA. Inference is like "doing the work"—users ask questions to ChatGPT, agents book flights, and code tools generate functions. The model has already been trained, and what's needed is to output results at the lowest cost and fastest speed.
Jalapeño is an inference-specific ASIC (custom chip), not a training chip. To use Privorotsky's analogy: GPUs are Swiss Army knives, while ASICs are chef's knives—the latter excel at specific tasks but cannot replace the former's versatility.
This means that OpenAI will continue to purchase NVIDIA GPUs for model training and research for the foreseeable future. Privorotsky explicitly stated that Jalapeño will not cause a "precipitous drop in NVIDIA demand," but that "the boundaries of discussion surrounding AI hardware are continuing to expand."
Hybrid transfer, not industry extinction
Privorotsky presents three core arguments in his analysis.
First, custom inference chips are no longer laboratory projects. The fact that a software company with no semiconductor history could collaborate with Broadcom to launch a chip with considerable performance in a very short period of time proves that the barrier to entry into this field has been significantly lowered.
Second, this does not mean that Nvidia's demand will decline this year. Training, research, and a large number of non-standardized workloads still require general-purpose GPUs, while the first-generation custom chips have limited shipments and long delivery cycles.
Third, looking at a longer timeframe, an increasing number of workloads in AI inference will migrate to dedicated chips. This is a "hybrid transfer"—it changes the distribution of profits, rather than the scale of the industry.
Goldman Sachs also pointed out that this event is not a negative for the memory industry. A 128-chip rack equipped with 27.5TB HBM4 means that OpenAI will pay higher fees per unit effective watt to memory suppliers such as SK Hynix, Samsung, and Micron. Migrating inference to custom chips will not reduce the demand for high-bandwidth memory; on the contrary, it may exacerbate scarcity.
AI-Assisted Chip Design: A Truly Underrated Narrative
At the end of his analysis, Privorotsky raised a question that has been largely overlooked by the market: How was a company with no history of chip design able to complete this chip in a record time?
Goldman Sachs' answer is: reinforcement learning and AI-assisted design cycle—using models to explore circuit schemes, compressing the design-measurement-verification cycle, and then using AI to generate software code that runs on the new chip.
This "recursive" logic means that AI is participating in the design of the machines that run AI. If this design cycle continues to be effective, the moat built around the software ecosystem by general-purpose GPUs will gradually thin in stable inference scenarios.
For the industry as a whole, compilers, kernel generators, and "models for writing code for chips" are evolving from peripheral tools into strategic infrastructure.
The re-stratification of ecosystems: who benefits and who is stressed
Nvidia: Short-term prospects are secure, but long-term structural pressures loom. Its training clusters, CUDA ecosystem, network connectivity, and the flexibility to "run new models on Tuesdays" still constitute a real competitive advantage. However, if the high-frequency daily workload of inference continues to migrate to ASICs, Nvidia will maintain its "gold rush" but gradually lose some of its "annuity." This is a multi-year issue concerning valuation multiples, not a shipment crisis in 2026.
Broadcom and custom chip foundries: becoming the "arms dealer" for every company that wants to develop its own chips but doesn't want to become TSMC. Custom inference is a service and networking business, and more Jalapeños mean more packaging, SerDes, and rack-mount manufacturing needs.
Model Labs: Jalapeño is both a profit weapon and a product weapon. Labs with their own inference chips can lower prices, increase speed limits, or allow agents to perform more steps before bills spiral out of control. Labs that only rent GPUs are renting someone else's cost structure; labs with inference chips can redefine the product itself. Google built TPUs for this, Amazon built Trainium, and OpenAI has just joined the race with its publicly available results.
Users and Pricing: Decreasing inference costs typically lead to two simultaneous results—lower prices for the base layer and higher prices for the front-end layer due to larger, more agent-like models. The decrease in total cost of ownership per token forms the digital foundation for simultaneously achieving a "$20 chatbot" and a "$2000 monthly agent service."
Risk Warning and DisclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.