Using AI routers to save tokens becomes a new trend? Giants scramble to lay out plans, enterprise AI costs drop by up to 97%
Against the backdrop of continuously rising AI usage costs, "model router" technology is rapidly penetrating the enterprise market.
The core logic is not complicated: automatically match the right, not the most expensive, model for each task, drastically reducing token expenditure without significantly affecting performance. This is prompting enterprises to shift from "defaulting to the strongest model" to "layered scheduling according to task."
At the industry level, model routers are making a leap from tool-type products to infrastructure-level components. Both tech giants and AI platform companies are accelerating the development of this capability, embedding it deeply into AI workflows and API gateways.
A deeper underlying signal is: the focus of AI competition is shifting from "model capability" to "scheduling and cost optimization capability"—whoever excels at fine-tuned allocation will be more likely to gain an advantage in the next stage.
From "default strongest" to "layered on demand"
The essence of the model router is cost-capability matching and stratification for AI tasks.
In actual enterprise usage, a large number of tasks do not require the strongest model. For example, scenarios such as email summarization, document abstraction, and information retrieval can be handled by smaller, cheaper models, while complex reasoning tasks are assigned to frontier models.
This approach became even more widespread after OpenAI released the GPT-5 series: systems automatically switch between different models according to task complexity. Subsequently, third-party routers that operate across models and providers emerged, enabling enterprises to dynamically schedule between models like OpenAI, Google, and Anthropic.
The multi-model collaborative system launched previously by Japanese AI lab Sakana AI shows that this routing mechanism already features certain "expert division of labor": mathematical problems tend to call OpenAI models, while scientific problems are more often routed to Google Gemini.
Cost reduction practices by leading enterprises: inference cost reductions of up to 97%
On the commercialization front, model routers are already leading to quantifiable cost reductions.
Palantir Technologies' routing tool, the Evolve AI routing system, not only handles model selection but also optimizes prompts and avoids duplicate calls. The company disclosed that in certain cases, by switching tasks from stronger models to lightweight ones, inference costs dropped by as much as 97%.
Construction company McCarthy Building also stated that its AI token usage dropped approximately 60% year-on-year, mainly due to model scheduling optimization.
Meanwhile, Databricks’ Unity AI Gateway has already been widely used internally. CEO Ali Ghodsi candidly pointed out that these tools are popular because enterprises are "burning through their AI budgets at too fast a rate."
On the security and network side, companies like Palo Alto Networks are also reducing AI call costs through model switching strategies, with routing capabilities gradually becoming a default module in enterprise AI architectures.
Capital and product resonance, model router sector accelerates formation
The capital market is also quickly following this trend.
Model routing platform OpenRouter completed $120 million in financing in April, becoming one of the most watched startups in this sector. Its core product, "automatic router," allows users to set a cost-quality preference range (0-10), and the system dynamically selects models based on this.
Data shows that about one-third of requests in its routing strategy are assigned to Google’s lower-cost models, while only about 10% are sent to OpenAI’s stronger models, clearly optimizing for cost stratification. OpenRouter’s underlying system integrates routing technology providers like Not Diamond, and supports cross-cloud providers to optimize latency and price structure.
Additionally, AI programming company Cognition has also launched its own routing system, which performs close to frontier models in programming benchmarks but lowers costs by about 35%.
Risk Statement and DisclaimerThe market involves risks; investment should be approached cautiously. This article does not constitute personal investment advice and does not take into account individual users' unique investment goals, financial situations, or needs. Users should consider whether any opinions, views, or conclusions in this article fit their specific circumstances. Investments made based on this content are your own responsibility.