Computing power shortage: Google quietly imposes Gemini usage limits on Meta
```
The supply-demand tension in AI infrastructure is intensifying among the world’s top tech companies. According to informed sources, Google informed Meta around March this year that it could not meet all of its Gemini computing power needs and imposed usage caps on the social media giant—even the world’s largest AI service providers are struggling to cope with the surging demand for computing power.
According to the UK’s Financial Times, these restrictions have not yet been lifted and have disrupted and delayed several internal AI projects at Meta. As a result, Meta has asked employees to improve the efficiency of AI compute usage and to carefully manage internal AI tokens. Both Google and Meta declined to comment.
This situation has forced Google to accelerate its expansion. Earlier this month, Google signed a $920 million per month compute capacity rental agreement with Elon Musk’s SpaceX. Google CEO Sundar Pichai admitted in this year’s first-quarter earnings call: “Recently we have indeed faced constraints in computing power; if we could meet demand, cloud business revenue would be higher.”
Meta is not alone. Several insiders said other Google enterprise customers have also been subject to varying degrees of restrictions, but Meta, due to its extremely large demand, has been affected the most. This incident reflects the explosive growth of AI inference workloads, which has become one of the biggest challenges facing the entire industry.
Computing Power Bottlenecks Under Pressure, Large Clients Hit Hardest
Despite major tech companies investing tens of billions of dollars in chips, data centers, and power supplies, the supply of AI computing capacity is still struggling to keep up with the pace of demand growth.
Google’s cloud business revenue exceeded $20 billion for the first time in Q1, and the backlog of signed but undelivered cloud contracts nearly doubled quarter on quarter, surpassing $460 billion. Pichai made it clear that computing power constraints will persist in the near term.
Against this backdrop, the blow to Meta is particularly pronounced. According to sources, it is precisely the strong demand from large clients such as Meta that is driving Google to quickly seek external sources of compute power. As enterprises deploy chatbots, programming assistants, and AI agents on a large scale, inference workloads—computing power consumed when models execute tasks in real applications after training—are becoming the industry’s core bottleneck.
Meta’s Internal Projects Thwarted, Accelerates Shift to In-House Models
Meta widely uses Gemini internally, covering platform content moderation (such as detecting scams and removing harmful content), customer service and advertising support chatbots, as well as some internal workflows and code development, with Anthropic’s Claude and other models used in conjunction.
According to informed sources, Meta initially chose Gemini because its performance was superior to the company’s in-house open-source Llama model. However, as compute restrictions tightened, Meta is accelerating its migration to its own models. Multiple sources say Meta has recently begun prioritizing the rollout of its newly launched Muse Spark model, believed to be comparable in performance to Gemini and to help reduce dependence on external models.
Meta CEO Mark Zuckerberg has continued to increase investment in AI talent and infrastructure, aiming to build what he calls “personal super intelligence.” Unlike Google, Meta doesn’t have a cloud business and is speeding up the establishment of its own data center system, pledging to invest $60 billion in the US by 2028.
Google Expands via SpaceX, Industry Seeks Solutions
Facing compute pressure, Google this month signed a $920 million per month compute capacity rental agreement with SpaceX to make up for infrastructure shortages. AI lab Anthropic also reached a similar agreement with SpaceX last month.
Google’s restrictions on Meta offer a rare window into the real pressures faced by the world’s top AI service providers in compute allocation. Currently, the entire AI industry’s infrastructure bottleneck is shifting from the training side to the inference side, and resolving the supply-demand imbalance still depends on the realization of a new round of large-scale capital investments.
Risk Warning and DisclaimerThe market has risks, and investment should be done cautiously. This article does not constitute personal investment advice, nor does it take into account any individual user’s specific investment objectives, financial situation, or needs. Users should consider whether any opinions, views, or conclusions in this article fit their specific circumstances. You are solely responsible for any decisions based on this article. ```