Ming-Chi Kuo: Nvidia restarts Rubin CPX project, an inference pre-filling optimization chip.
Nvidia has quietly restarted a chip project that was once thought to have been abandoned by the market, targeting the prefill stage, which has high AI inference costs.
According to a recent industry survey by renowned analyst Ming-Chi Kuo, Nvidia has restarted its AI inference pre-filled accelerated GPU project, "Rubin CPX," and made significant adjustments to the product design. Based on current plans, the new version of Rubin CPX is expected to begin production in the first quarter of 2027.
The restart of this project sends a clear signal: NVIDIA is ramping up its efforts to address the performance and cost bottlenecks of pre-filling in the AI inference stage. As the context length of large models continues to increase, the pre-filling stage needs to process a large amount of input information and generate a key-value cache, and its efficiency directly affects the overall cost and deployment economy of AI inference.
According to Ming-Chi Kuo, more than 50% of AI inference workloads currently come from input context processing and key-value cache construction.
The chip specifications have been significantly adjusted, and the computing power is close to that of standard Rubin.
The core specifications of the new Rubin CPX have changed significantly compared to the previous version.
In terms of computing power and power consumption, the new CPX's performance is close to that of the standard Rubin GPU, with a maximum power consumption of 2300 watts per card. Regarding memory, the new CPX uses 168GB of HBM4 memory, a significant upgrade from the previous 128GB GDDR7, but still less than the standard Rubin GPU's 288GB HBM4.
According to Ming-Chi Kuo, the total HBM4 capacity of the 8-card CPX compute tray is approximately 1.34TB, which can cover most long-context pre-filled workloads and their corresponding KV cache requirements. This means that Rubin CPX is not simply pursuing higher general-purpose computing power, but rather has redesigned memory capacity and pre-filling efficiency for long-context scenarios.
Shifting from shared racks to independent deployments allows for more flexible configuration.
Rack architecture is also a key focus of this adjustment.
In the previous version, CPX planned to share racks with Rubin GPUs; the new version uses a separate MGX ETL rack, allowing customers to expand CPX computing power individually according to their actual needs.
Specifically, customers can choose deployment sizes of 64, 128, 192, or 256 CPX GPUs. Each rack module consists of 64 CPX GPUs, including 8 compute trays, each tray carrying 8 CPX GPUs, and a switch tray.
Modular design means that customers do not need to purchase in fixed large-scale Rubin racks, and can flexibly increase CPX computing power according to the actual needs of pre-filled load.
Layered interconnect design reduces pre-populated deployment costs
In terms of interconnect architecture, Rubin CPX adopts a layered design, making trade-offs between performance and cost.
Within a single compute tray, eight CPX GPUs are scaled up via NVLink. Each CPX GPU has an NVLink bandwidth of 1 to 1.5 TB/s, lower than the 3.6 TB/s of a standard Rubin GPU.
At the scale-out level between trays and within rack modules, CPX uses Spectrum-6 Ethernet all-copper L1 links; when connecting across rack modules, interconnection is completed through OSFP fiber links via Spectrum-6 switches in each module.
Compared to solutions that rely entirely on high-bandwidth GPU interconnects, this layered architecture places greater emphasis on cost optimization for pre-filled scenarios.
Working in conjunction with Rubin GPUs to target long-context inference
Rubin CPX is not a standalone general-purpose GPU, but rather a pre-filled accelerator used in conjunction with Vera Rubin NVL72.
According to Ming-Chi Kuo, NVIDIA recommends deploying CPX and Rubin GPUs in a 1:1 ratio: CPX is responsible for the pre-filling stage calculations and generating the KV Cache, which is then transferred to the Rubin GPU via Ethernet RDMA for subsequent decoding.
From a product positioning perspective, the core of Rubin CPX is not to replace the standard Rubin GPU, but to further break down the AI inference process, allowing different GPUs to undertake pre-filling and decoding tasks respectively.
Ming-Chi Kuo positions it as "the best cost-effective solution for long context pre-filling." As the context window of AI models continues to expand, the computational load and KV cache size in the pre-filling stage continue to grow. NVIDIA's restart of the CPX project may be an attempt to further reduce the cost of long context inference through dedicated hardware and heterogeneous deployment.
Risk Warning and DisclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.