Meta's customized AMD MI450 chip has its computing power halved and memory drastically reduced. SemiAnalysis: This is a tragedy of Meta's company culture.

Meta's customized AMD MI450 chip has its computing power halved and memory drastically reduced. SemiAnalysis: This is a tragedy of Meta's company culture.

```

Meta's series of missteps in the area of AI infrastructure are exposing deep organizational cultural problems at the cost of billions of dollars.

According to a recent disclosure by the chip industry research firm SemiAnalysis, Meta is requesting AMD to customize a significantly downgraded version of the MI450X chip—halving compute units, reducing the HBM memory stack from 12 layers to 8, and shrinking memory capacity by nearly two-thirds. SemiAnalysis directly labeled this decision as "catastrophic," and publicly urged AMD to bypass Meta’s infrastructure team and directly liaise with Meta’s super-intelligence lab TBD Lab, advocating procurement of the standard MI450X.

This incident is not isolated. SemiAnalysis points out that from the over-$2.5-billion Rivos acquisition, to the H100 custom server Grand Teton, and the GB200 custom solution Ariel, Meta’s infrastructure team has consistently shown traits of over-engineering, a lack of hardware-software co-design, and short-term political motivations overruling long-term technical rationality. "Meta’s infrastructure team needs a cultural reset," SemiAnalysis writes.

AMD’s Top Chip “Castrated,” GenAI Performance Severely Damaged

AMD MI450X is currently one of the most aggressively engineered GPUs on the market: using 2nm process, hybrid bonding packaging, 12-layer HBM4 memory stack, and the largest CoWoS mask size in the industry, representing the ceiling for packaging and storage density today.

However, according to SemiAnalysis, the custom version ordered by Meta will cut its compute silicon area by half, reduce HBM stacks from 12-Hi to 8-Hi, compressing both compute power and memory bandwidth, its core indicators. Meta’s infrastructure team claims this configuration is designed for recommendation system (RecSys) workloads, aiming to improve the compute ratio between CPU and GPU.

The issue is this decision was made before the formation of TBD Lab, which is Meta’s core large model research team and has no interest in this chip. SemiAnalysis clearly states that compared to Nvidia’s Vera Rubin, the downgraded MI450 holds no appeal for TBD Lab, "TBD will heavily favor Rubin." This means AMD’s shipment volumes to Meta will suffer severely because of this decision.

SemiAnalysis, unusually, directly called out to AMD in its report: "AMD needs to take a stand, work directly with the TBD team to ensure they get the standard MI450, rather than this castrated version that is useless for GenAI."

Rivos Acquisition: Over $2.5 Billion for a Mess

Another typical case of Meta’s infrastructure decision errors is the over-$2.5-billion Rivos acquisition in 2024.

According to SemiAnalysis, few within Meta’s chip department truly understand the strategic logic behind this acquisition, and those who initially drove the deal have since become silent. The mainstream view is: Meta has abundant funds, the custom chip track is heating up, and since it was already licensing Rivos IP, the management thought they might as well acquire it outright.

However, the transaction structure itself contained hidden risks. The founding team of Rivos insisted on selling as a complete package, forcing Meta to buy the whole company and later cut departments it didn’t need. The report cites former Meta chip employees saying that the acquisition was led by Meta’s silicon head Yee Jiun Song, who encountered internal opposition, but lost interest after the deal was done. The current chip team managers treat Rivos engineers as "free labor," assigning them to their own teams, and the original structure quickly dissolved.

From a technical perspective, the core value of the acquisition—the SIMT architecture GPU IP from Rivos—vanished when the Olympus chip project, intended to use it, was canceled. The replacement "Phoebe" project is expected to tape out as early as 2028, but SemiAnalysis notes that there’s little optimism internally for its successful implementation.

Staff losses are also alarming. According to SemiAnalysis, about 30% of Rivos employees left in the recent round of layoffs, co-founder Mark Hayter has departed, and several former Rivos staff moved to Nuvacore, a chip startup founded by Gerard Williams, after their first batch of RSUs vested in May. SemiAnalysis also disclosed that Rivos CEO and co-founder Puneet Kumar may leave one or two years after his Meta shares fully vest.

Grand Teton and Ariel: Repeating the Costly Design Mistakes

Meta’s obsession with customization is nothing new. Its H100 custom server Grand Teton, based on the standard HGX server, added an extra switch tray, with four Broadcom PCIe switches, 16 SSDs, and eight network cards, to offer more direct-attached storage per server for saving training checkpoints.

Yet after production, the model team used this storage far less than expected and the design was ultimately abandoned. SemiAnalysis notes this is a typical example of Meta’s hardware and software teams lacking co-design: the infrastructure team paid extra for materials and power for features that were not fully used, and increased dependence on Broadcom, contradicting Meta’s original intent to reduce dependence on Nvidia’s networking gear.

With the Blackwell era, Meta’s custom plan "Ariel" follows similar logic. The standard GB200 spec pairs each Grace CPU with two B200 GPUs, but Ariel changed this to a one-to-one match, halving the GPU count. The rationale is again to boost the CPU ratio for RecSys workloads.

According to SemiAnalysis's calculations, the total cost of ownership (TCO) of the Ariel NVL36x2 plan is 14% higher than the standard GB200 NVL72, and the extra spending yields more CPU and DRAM resources—ironically, exactly what the large model team doesn’t need. Because of inter-rack connections used to compensate for fewer GPUs, network latency and reliability problems arise. Meanwhile, the originally considered unstable NVL72 backplane has matured, and Meta’s risk assessment proved to be wrong in hindsight.

SemiAnalysis says that Meta’s GB300 server has reverted to standard configuration—essentially a denial of the Ariel plan.

Cultural Roots: Short-term Evaluation Overwhelms Long-term Strategy

SemiAnalysis attributes these problems to the culture of Meta’s infrastructure team.

The firm notes that Meta’s six-month performance evaluation cycle eliminates the bottom 10–15% employees each round, causing the team to favor short-term visible results over long-term technical planning. Some managers are fond of advancing highly visible but quickly deliverable projects, then swiftly pivot after delivery—internally called "window washing." Meanwhile, few openly challenge decisions made by superiors, making it difficult to correct mistakes promptly.

The supply chain team has limited say in engineering decisions, further amplifying these issues. SemiAnalysis points out that some suppliers, due to Meta’s frequent direction changes, have reduced the priority of its new projects, instead safeguarding Amazon or Google’s needs first.

SemiAnalysis likens this phenomenon to the growth path of Meta’s Reality Labs: before mass layoffs, billions were poured into massive engineering teams and R&D projects. Now, as Meta sells computing power to external customers, the cost of this internal culture will be directly exposed to the market.

Risk Warning and DisclaimerThe market is risky; investment requires caution. This article does not constitute individual investment advice, nor does it consider the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their particular situation. Investing accordingly is at your own risk. ```