When AI starts “teaching itself”: Is Opus 5.5 the first large model trained with RSI?

When AI starts “teaching itself”: Is Opus 5.5 the first large model trained with RSI?

Anthropic's latest release, Claude Opus 5.5, boasts a significant performance boost while reducing operating costs by 40%, an unconventional combination that has drawn close attention from the industry to its training methods.

On the social platform X, Google engineer Patrick C Toulme made a speculation: Opus 5.5 may have been distilled from the more powerful "teacher model" Model 2 within Anthropic, and may be the first product model to embody the "Recursive Self-Improvement" (RSI) path.

There is currently no publicly available evidence to directly confirm this inference, but Anthropic has publicly confirmed the existence of the Model 2, and the company's RSI-related articles show that AI's involvement in AI research and development is rapidly deepening, which gives the above speculation some background.

For investors and industry observers, what's truly noteworthy about this event is that it connects three previously relatively separate trends: continuous iteration of stronger internal models, deep integration of AI into the AI R&D process, and the simultaneous emergence of smaller and cheaper product models .

If this path holds true, improvements in model capabilities and reductions in inference costs could potentially occur simultaneously.

Why is Opus 5.5 worth paying attention to?

According to Wall Street News , Anthropic positions the Claude Opus 5.5 as the first release in its 5.5 series. Its performance is on par with the Claude Fable 5.1 in most tasks, but its operating cost is 40% lower than the Opus 5 and its output speed is more than 30% higher.

The price reductions are particularly significant . Input and output token prices are $4 and $20 per million, respectively, a 20% decrease from Opus 5; while cache reads, which account for the majority of costs in proxy tasks and programming scenarios, have dropped from $0.50 to $0.20 per million, a 60% decrease.

In terms of actual usage , early testers completed the migration of 680,000 lines of code in less than a day, while similar work previously required several weeks for the engineering team. In another test, Opus 5.5 achieved a 39/40 success rate in optimizing the overall page load time of a web application, while Opus 5 exhibited side effects from altering application behavior in the same task.

In internal tests rewriting HAProxy from C to Rust, Opus 5.5 took 9.5 hours, more than 20% shorter than Fable 5.1's 12 hours, resulting in a 51% cost saving.

Anthropic stated that the efficiency advantage is also reflected in the token consumption level – Opus 5.5 not only has a lower price per token, but also requires fewer tokens to complete the same task, resulting in a combined 40% reduction in overall costs.

What exactly is the "teacher model" that Patrick is referring to?

Patrick C. Toulme wrote on X:

Opus 5.5 was clearly trained from a larger teacher model, most likely Model 2 Mythos. Opus 5.5 is the first model trained with RSI and simultaneously distilled from an internal teacher model. Its smaller size and lower cost are a direct result of teacher model distillation.

The core technical concept involved here is "model distillation": a more capable teacher model generates high-quality training signals, and these signals are then used to train a smaller student model, thereby significantly reducing inference costs while retaining high capabilities.

Anthropic has publicly confirmed the existence of a more powerful Model 2 than Mythos 5, which is extensively used for code generation, data generation, and proxy tasks. However, there is currently no publicly available material directly proving that Opus 5.5 is distilled from Model 2. Anthropic's release documentation also makes no mention of specific training methods.

It is worth noting that Anthropic lists "distillation attacks" as a separate threat category in its security terms—that is, attackers extract capabilities from the model in bulk through a large number of fake accounts, and for this purpose, it introduced anti-distillation mechanisms such as "preserving the inference chain" in Opus 5.5.

This detail indirectly illustrates that distillation technology itself holds a very important position in Anthropic's technological system.

AI begins to participate in AI research and development.

Regardless of the specific training path of Opus 5.5, Anthropic's own RSI-related articles provide a more macro-level picture: the depth of AI's involvement in AI research and development is accelerating at a quantifiable pace.

As of May 2026, over 80% of the code in the Anthropic codebase was written by Claude, compared to a single-digit percentage before the Claude Code research preview was released in February 2025.

In terms of engineering output, Anthropic engineers have increased their average daily merged code volume by about eight times compared to 2024.

The Anthropic article also points out that in April 2026, Claude reduced a certain type of API error by 1,000 times in about 800 hours, while the engineers who led the work estimated that it would take about four years to complete the same task manually.

In terms of research judgment, Anthropic designed an internal test: in collaborative conversations between researchers and Claude, they identified key points where human researchers chose "suboptimal" directions, and then compared the merits of different versions of Claude and human judgments at these points.

The results showed that the Opus 4.5 model in November 2025 provided better next-step suggestions than humans in 51% of cases, while the Mythos Preview in April 2026 saw this proportion rise to 64%.

However, the article also clearly points out the current boundaries:

Claude still lags far behind humans in choosing which research questions to pursue. This is precisely the distance between today's AI and future systems capable of autonomously designing the next generation of AI.

Anthropic also emphasized that fully recursive self-improvement has not yet been achieved, nor is it inevitable.

However, analysts believe that while past model competition largely boiled down to "who has more computing power," a new path is emerging: using stronger models to help train and improve cheaper models. If this path holds true, improvements in model capabilities and reductions in inference costs could occur simultaneously.

Risk Warning and DisclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.