Is the memory bottleneck still unresolved? The AI chip industry is accelerating its performance by "decoupling memory".

Is the memory bottleneck still unresolved? The AI chip industry is accelerating its performance by "decoupling memory".

The memory wall is becoming a core obstacle to the expansion of AI infrastructure, and the semiconductor industry is using architectural restructuring as a breakthrough.

According to a recent report by Bank of America Securities research team, on the second day of the 2026 AI Infrastructure Summit, major vendors and cloud computing giants such as Micron, Samsung, SK Hynix, Broadcom, Marvell, Intel, Qualcomm, OpenAI, AWS, and Google made appearances, with the core topic focusing on strategies to address the "memory wall" problem. Participants generally agreed that the layered disaggregation of memory and storage has become the most widely accepted technical approach.

This trend has a direct impact on the semiconductor investment landscape: efficiency metrics are replacing computing power as the new industry benchmark, shifting from raw computing power (FLOPs) to tokens per watt (tokens/W) or tokens per dollar (tokens/$); at the same time, interconnect architecture is accelerating towards openness, and the demand for customized chips is rising in tandem, thus attracting market attention to the product portfolios of companies such as Broadcom, Marvell, and Astera Labs.

Memory wall: The core bottleneck of AI expansion

The growth rate of memory bandwidth and capacity is far behind the expansion of model size—this is a core judgment that this summit continues from the Hot Chips conference in August this year.

The size of Transformer models grows approximately 240 times every two years, while memory bandwidth and capacity only increase by about 2 times every two years, and the gap between the two is widening. This structural imbalance makes the traditional GPU-centric architecture increasingly unsustainable, driving the industry to shift towards specialized, tiered memory solutions.

At the summit, Qualcomm showcased its "HBC" (High Bandwidth Computing) solution, which achieves approximately 200 times the capacity/power ratio of SRAM and approximately 6 times the bandwidth/power ratio of HBM by directly stacking LPDDR memory on top of the computing chip. Samsung's zHBM also employs a similar 3D DRAM stacking approach.

SK Hynix highlighted three dedicated memory tiers: first, PIM (In-Memory Compute) for memory-intensive workloads, which offers 288 times the capacity per rack compared to SRAM; second, HBF (High Bandwidth Flash Memory) for long-context scenarios, with a capacity approximately 10 times that of HBM; and third, the SALT-KV software solution, which enables cross-tier temperature-sensing KV cache scheduling.

Efficiency First: The New Laws of AI Scalability

The summit sent a clear industry signal: the era of stacked computing power is giving way to a new paradigm driven by efficiency.

A Bank of America Securities report points out that as AI applications evolve towards agency AI, industry evaluation systems are shifting from pursuing raw FLOPs to focusing on efficiency metrics such as tokens/W and tokens/$. This shift directly drives the demand for memory diversification and layered decoupling—different types of workloads require different memory tiers with different characteristics, rather than relying on a single high-performance memory solution.

This has a substantial impact on the procurement logic of data center operators and the product strategies of chip manufacturers, and the flexibility and energy efficiency of memory architecture will become new dimensions of competition.

The parallel advancement of interconnectivity, openness, and customization

At the network interconnection level, Ethernet has basically established its dominance in the field of horizontal scaling (scale-out), and openness and interoperability are the key to its success.

Broadcom is pushing Ethernet to every level of the interconnect architecture: in terms of horizontal scaling, Tomahawk 6 (102.4T) has entered high-volume production and is deployed in hyperscale cloud vendors; in terms of vertical scaling, Thor Ultra NIC and ESUN solutions are ready; and in terms of cross-domain interconnect, Jericho 4 takes the lead—all using an open architecture and not relying on vertical integration.

Meanwhile, the demand for customization is also rising. Marvell offers end-to-end solutions covering custom XPU attach, all-optical interconnect kits, and multi-protocol support (UALink, NVLink Fusion, ESUN); Astera Labs, on the other hand, is based on the general PCIe protocol and is deploying in the interconnect and memory fields through dedicated chips such as Scorpio and Leo.

Mass production speed and reliability: equally important as chip design

At the summit, cutting-edge laboratories and data center operators conveyed a consistent message: the speed of rack mass production and system reliability are now as important as chip design itself.

AWS has significantly shortened the traditional 6 to 9-month post-silicon testing and stabilization cycle for chips by matching testing capacity with installation scale and coordinating manufacturing processes with deployment locations, thus significantly reducing the time lag between chip release and actual data center deployment.

Software co-design is also listed by hyperscale cloud vendors as a key tool to accelerate deployment. Capabilities such as native PyTorch support and Hugging Face portability, which can be achieved with minimal code changes, are considered important means to shorten the cycle from chip to application deployment.

Bank of America Securities research team believes that a rapid product iteration cycle of more than one year is crucial for token economics and market competitiveness. The realization of this cycle no longer depends solely on chip design capabilities, but also on system-level engineering integration and hardware-software synergy.

~~~~~~~~~~~~~~~~~~~~~~~

The above content is from Zhuifeng Trading Platform .

For more detailed analysis, including real-time updates and firsthand research, please join the [ Trading Channel Annual Membership ].

Risk warning and disclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.