U.S. tech companies are quietly turning to Chinese AI models, with Coinbase leading the way in adopting GLM and Kimi.
```
American technology companies are quietly incorporating Chinese open-source AI models into their production infrastructure. As the costs of top U.S. models continue to rise, companies like Coinbase are making Chinese open-source models the default option, significantly reducing AI expenses without restricting usage.
Coinbase CEO Brian Armstrong revealed in a post on X last Friday night that the company has set GLM 5.2 from Zhipu and Kimi 2.7 from Beijing Moonshot as default models for engineers through its internal LLM gateway. Armstrong stated that, after combining routing optimization and caching improvements, Coinbase's AI expenditure has been cut by "nearly half," while token usage continues to grow exponentially.
Cost advantage of Chinese open-source models comes to the forefront
Armstrong made it clear in his post that 91% of engineers had never reached their previous usage limits, so Coinbase did not choose to lower the limits or add consumption alerts, but instead switched to "cheaper default models."

GLM 5.2 is from Zhipu, Kimi 2.7 is from Beijing Moonshot, and both are open-source weighted models. Armstrong said these models are deployed for routine task scenarios, while engineers can still opt for cutting-edge models for complex planning tasks. His reasoning: using top models for execution-level tasks is often "overkill."
For code review, a multi-model parallel strategy is adopted, where different models cross-verify output to maintain quality standards.
Three-layer infrastructure restructuring drives cost reduction
Armstrong listed three key approaches.
The first is smart routing: in a custom scheduling framework, the system preprocesses prompts, combines cache hit rates and model pricing, and automatically assigns tasks to the most suitable and economical model. He said the ultimate goal is for AI, not humans, to make the model selection.
The second is aggressive caching: Coinbase requires all requests to have cache awareness and to reuse existing cache as much as possible. For example, with proper caching mechanisms in LibreChat, cache hit rates jumped from 5% to 60%.
The third is trimming context: Armstrong recommends starting new sessions when switching tasks, narrowing the file context range, and disconnecting unused tool integrations. He emphasized that the goal is not to reduce total token usage but to reduce "wasted tokens."
Efficiency first, not usage suppression
Armstrong characterizes this cost reduction as a prerequisite for scaling up AI adoption, not a restriction. He says engineers are still free to use any quantity of tokens and any model, but the company now visualizes usage data and ties it to business impact—"the more we spend, the greater the impact expected."
He did not disclose specific absolute spending figures. Structurally, achieving nearly half the expenditure reduction while usage grows exponentially means Coinbase has, to some extent, decoupled consumption from costs.
Armstrong concludes that this methodology is universally applicable—any company can use it to achieve sustainable expansion of AI usage without making cost the ceiling.
Risk Warning and DisclaimerThe market carries risks; investment requires caution. This article does not constitute personal investment advice, nor does it take into account the special investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article suit their specific circumstances. Investments made accordingly are at the user's own responsibility. ```