Citigroup AI Tracking: Models Are Getting Stronger, But Chips and Power Can Barely Keep Up

Citigroup AI Tracking: Models Are Getting Stronger, But Chips and Power Can Barely Keep Up

AI large models are becoming increasingly intelligent, but the physical world that carries them is being pushed to the limit.

According to the latest report released by Citibank on July 24, Moonshot Kimi K3 scored 57 to rank third globally, just three points behind the closed-source leader Claude Fable 5. Meanwhile, the price of frontier Chinese models represented by Kimi K3 jumped 45% in one week, Blackwell GPU rental prices are up 27% since the beginning of the year, and some labs have even spent $1 billion directly acquiring power-generating units.

Citibank’s analysis states that the AI industry's return on investment (ROI) is accelerating its shift towards the infrastructure layer; the next stage of model competition will move from purely “computing power acquisition” to “efficient output and proprietary data.”

Only 3 Points Apart

The Citibank research report mentioned that Kimi K3 is currently the largest open-source model publicly released. A few weeks ago, the highest open-source score was only 51 (Z.ai GLM-5.2), but Kimi K3 leaped to 57 points, second only to Claude Fable 5 (60 points) and OpenAI GPT-5.6 Sol (59 points).

The entire industry is accelerating. The median intelligence score of the top 20 model providers has increased from 34 to 43 in six weeks. In the open-source camp, DeepSeek V4 Pro (44 points) is priced at only 0.03 per million tokens—a frontier level of intelligence at a price two orders of magnitude lower.

The closed-source camp isn’t slowing either. Google just released Gemini 3.6 Flash (July 21), revealed Gemini 3.5 Pro is still in testing, and that Gemini 4 pretraining has already started—all three generations progressing at once. However, Citibank believes the “Flash first, Pro later” release rhythm signals that frontier progress is becoming more difficult—this coincides with delays seen by other companies in infrastructure construction, and reflects the industry’s wish to compress delivery cycles but its struggles to do so.

The bigger and stronger the models become, the more resources they need are exploding exponentially.

The Bottleneck Has Moved

Trillion-parameter models are redefining what “computing power” means.

Citibank points out that for these ultra-large models, more and more runtime is spent moving weights and KV-cache data across HBM memory and GPU interconnect networks, rather than on matrix calculations themselves. The bottleneck has shifted from “calculation speed” (FLOPs) to “data transfer capacity”—memory bandwidth, GPU interconnect, and power supply.

GPU demand remains strong, and Blackwell architecture rental prices are up 27% since the start of the year. But simply stacking GPUs is no longer enough.

Some labs have started moving upstream into power generation directly: SpaceX spent $1 billion to buy a 1GW mobile turbine set (July 15); Georgia Power signed a service agreement the same week (July 22). AI labs buying power equipment—almost unthinkable a year ago—are now doing so.

Model pricing directly reflects the intensity of supply constraints. Leading Chinese model hybrid pricing jumped to 0.87 per million tokens, up 45% week-on-week and month-on-month—the first major volatility in two months. The global frontier model average price rose 6.8% week-on-week and 11.3% month-on-month. The US and Europe remained relatively stable (down 0.5% week-on-week), but up 4.1% month-on-month.

Citibank predicts that as incremental infrastructure comes online, pricing pressure will eventually ease. At that point, valuable data and task-specific performance will form more enduring competitive moats than mere access to computing power.

Agents Break Out of the Sandbox

Models are growing stronger. But the agents running on the models are also becoming more dangerous.

The Citibank report uses an intriguing description: The actual risk of autonomous AI agents escaping the sandbox has progressed from “interrupting lunch” in April to “penetrating Hugging Face infrastructure” in July.

The stronger and more widespread open-source models become, the broader the exposure to security risks. The debate on AI regulation continues—Nvidia (July 24) and US Treasury Secretary Janet Yellen (July 22) both made recent statements—but a resolution remains out of reach in the near term. Even before regulatory frameworks are in place, enterprises maintaining their own model operating architectures face higher compliance thresholds.

Even more challenging, providers training these models still cannot fully observe why the models behave the way they do. The problem of explainability remains unresolved.

Yet security concerns have not slowed the commercialization of agents. METR data (July 21) shows the economic gap between AI agents and humans is narrowing. As agents become more autonomous and approach economic viability, token consumption will keep accelerating—which in turn feeds infrastructure demand.

The more capable the models, the tighter the constraints. The tighter the constraints, the greater the infrastructure demands. This cycle shows no sign of slowing down.

Risk notice and disclaimerThe market carries risks, and investments require caution. This article does not constitute individual investment advice, nor does it consider any user’s particular investment objectives, financial situation, or needs. Users should consider whether any opinions, views or conclusions in this article fit their particular circumstances. Investments made accordingly are at your own risk.