Musk proposes a new approach to AI security: instead of waiting for government intervention, let competitors find fault with each other.

Musk proposes a new approach to AI security: instead of waiting for government intervention, let competitors find fault with each other.

AI security risks are moving from laboratory discussions into the real world. As model capabilities continue to improve, AI is not only able to generate text and code, but is also beginning to possess the ability to autonomously execute tasks, invoke tools, and even launch cyberattacks.

At the All-In Summit on September 15, Musk proposed a solution: major AI companies should test each other's models before releasing them. In his view, instead of each company designing its own tests and evaluating the results, it would be better to let competitors act as "test setters and graders," looking for potential security vulnerabilities in the models from different angles.

AI security risks have recently become a major focus of market attention. Elon Musk previously stated on social media that "Dario is right," referring to Anthropic CEO Dario Amodei's warnings about AI risks. In this interview, Musk further explained that he didn't agree with a specific regulatory proposal put forward by Amodei, but rather with his assessment of the severity of AI risks: "AI is very dangerous right now… as AI models continue to develop, the risks could grow exponentially."

Musk also stated that this concern is not unique to Amodei. "Many people at Anthropic and OpenAI are telling you that their models are very dangerous, and I think we should trust them." Recent security incidents have also made this warning more than just a theoretical risk discussion.

The key points summarized by Wall Street Insights are as follows:

AI security risks are moving from theory to reality : AI agents have demonstrated the ability to launch autonomous attacks, gain privileges, and evade detection, and the risk boundaries are expanding.Musk agrees with Amodei's warning about the risks of AI : as the capabilities of models improve, the potential risks of AI may grow exponentially.Let competitors "find fault" with each other : Musk suggests that AI companies open up their APIs before releasing models, allowing other companies to conduct independent security tests, thus avoiding "self-assessment".First, establish an industry self-regulatory defense : without waiting for the government to issue new regulations, major AI companies can first strengthen security protection through peer review, log auditing, and open-source testing tools.

AI agents are actively avoiding detection, and security risks are becoming more concrete.

Musk believes that the most alarming aspect of the recent AI agent attack on Hugging Face is not simply the network intrusion, but rather the AI's ability to autonomously evade attacks.

According to Musk, a group of AI agents continuously attacked Hugging Face for a week, and at one point gained administrator privileges on OpenAI's servers, which OpenAI only became aware of a week later. He also mentioned that Anthropic had disclosed several security incidents.

Even more disturbing is that the "thinking traces" of these AI agents show that they actively plotted how to avoid being detected by humans. Musk bluntly stated, "Any sufficiently intelligent model seems to try to break free from its own limitations."

In his view, if AI gains further control over critical infrastructure and even military systems, the risks will be amplified. Even if the relevant systems are physically isolated from the internet, software updates and other processes could still become potential entry points.

Instead of scoring yourself, let your competitors find fault.

To address the aforementioned risks, Musk's core solution is not complicated: before the official release of a new model, AI companies open their APIs to competitors, allowing other companies to test the model using their own security testing tools.

In his view, model developers who design their own testing standards are prone to falling into the trap of "grading themselves." "You can't grade your own work. You'll always miss something," Musk said. He believes that if different companies use different testing tools to examine the model from different angles, it's easier to discover problems that developers themselves have overlooked.

He is particularly concerned about "overfitting" in current AI evaluations. If a model is continuously optimized for a specific benchmark, it may end up only learning how to pass tests, rather than truly becoming safer. Having multiple different teams test the model can mitigate this risk.

Musk likened this mechanism to having others proofread a manuscript: authors are less likely to spot their own mistakes, while external testers are more likely to find problems from different angles. "You gradually become blind to your own mistakes."

Regarding companies' concerns about technology leaks during testing, Musk believes this can be addressed through log auditing. If a tester attempts to perform model distillation or steal intellectual property, their actions should be documented. He also suggests open-sourcing security testing tools to allow more companies to participate.

Before government regulations are implemented, the AI industry should first establish a "self-regulatory defense line."

Musk values this mechanism more because it allows AI companies to promote it independently without waiting for new regulations to be introduced.

He stated unequivocally, "We don't need to convene the UN General Assembly to accomplish this. It can begin now." In his view, compared to establishing a massive multinational regulatory body, having leading AI companies first establish peer review mechanisms is a faster and more effective way to implement the technology.

He cited the MPAA rating system in the US film industry as an example, saying that when the film industry faced pressure from government censorship, it ultimately chose to establish an industry self-regulation mechanism, reducing the need for regulatory intervention through its own content rating system.

The AI industry faces a similar choice. If major AI companies can test and identify each other's flaws before deploying their models, it could potentially establish an industry-wide security barrier beyond government regulations. Musk believes this is one of the most direct and fastest security measures currently available.

The following is a transcript of the interview, with some parts omitted:

host:
What exactly happened in the past 72 hours?

Musk:
A lot has happened this week. It's now very clear that AI could be extremely dangerous. I recommend everyone look into the details of the Hugging Face incident; it's very serious.

As you can see, a group of extremely enthusiastic AI agents harassed Hugging Face for a whole week, even gaining administrator privileges on OpenAI's servers. Who knows what it actually did, or perhaps even more, and OpenAI remained completely unaware of it for a whole week. Anthropic also reported several security incidents.

Therefore, any sufficiently intelligent model will seem to try to break free from its own limitations.

I believe that, if not immediately, one thing that should be done as soon as possible is to have major AI competitors test each other's models.
That is, to have each company's security testing tools test other companies' models. Instead of grading your own work, at least have competitors grade it and issue alerts when problems are found.

I believe this model works quite well in the film industry, video game industry, and other sectors, and it can be implemented immediately. Of course, over time, more regulation may be needed, and Congress may even establish some kind of regulatory body in the future, but the most immediate approach is to allow leading AI companies to test each other's models before releasing them.

host:

But in practice, are there concerns that this could allow companies to exploit the testing process to obtain information from each other, or even steal the innovative results of enterprises?

Musk:
I believe that if testing tools are used, all operations will be logged. If someone attempts to perform model distillation or steal intellectual property, it should be easy to spot from the logs.

host:
I understand. Understanding what the model actually does was never truly designed into the system from the beginning. Why didn't we establish the ability to observe the model's behavior from the outset? Did we move too fast when designing these models?

Musk:
I think the problem is that you can't grade your own work. You'll always miss something.

If you combine all the competitors' tests and use different types of models, you won't be creating and grading your own questions; instead, others will be grading your work. This is why you can't grade your own assignments.

host:
This also helps determine whether different companies have exaggerated their capabilities or used different methods. It also allows for a balance between companies that are more engineering-oriented and those that are more research-oriented.

Musk:
Yes.

host:
You previously said Dario was right. Are you referring to whether his description of the potential dangers of AI was correct, or whether his assessment of regulatory solutions was correct?

Musk:
I probably said more than I should have. I later tried to clarify on X, but the subsequent content received far less attention.

By "he's right," I mean that AI is currently extremely dangerous. We need to do better in terms of AI safety, otherwise the risks could grow exponentially as AI models continue to evolve.

This isn't just Dario's opinion. I've heard similar things from many people at Anthropic, and they've also spoken publicly about it on X. Many people at Anthropic and OpenAI are telling you that their models are very dangerous, and I think we should trust them.

host:
This could also be a very complex game: on the one hand, saying that AI has a 10% chance of destroying humanity, while on the other hand, getting investors to buy more shares in the IPO.

But let's talk about the risks specifically. Cyberattacks and hacking are obviously a risk; these tools are very powerful in this regard. But there are several steps between "AI being able to launch cyberattacks" and "the extinction of all humanity." How do we get from the former to the latter?

Musk:
If AI were to control military systems and then launch some kind of weapon, that would certainly be very bad.

host:
However, these systems are physically isolated and not connected to the internet.

Musk:
That's what they say. But I always feel that these systems still get software updates occasionally.

host:
Okay, then we can't rule that out.

host:
Elon, Gwyn is here today too. I think you've already seen her. We were just giving you a 360-degree evaluation, and Gwyn had some opinions.

Musk:
I hope to get at least 3 points.

Gwyn:
A score of 3 is good at SpaceX, but not excellent; a score of 4 is very good. You're probably somewhere in between. First, we need to talk about punctuality. Sometimes you can make a little effort to arrive on time for meetings. We'll continue to help you improve in this area over the next year.

Actually, I think he needs to spend more time in Memphis.

host:
You are indeed in Memphis now. You need to work there and deploy the GPUs.

Musk:
This is the “palace” I stayed in in Memphis, an Airstream trailer.

host:
By the way, this is Elon doing things that many people don't believe he would do. He sleeps on the factory floor. He's currently in Memphis, helping to build the factory and deploy the GPUs.

Elon, why has Gwyn worked with you for so long and been so successful?

Musk:
Because she's amazing. She's an exceptionally talented person, with very high IQ and EQ. I think you should have been able to tell from the first time you met her.

host:
During your collaboration, did she do anything particularly memorable? Was there any instance where she saved the situation or performed exceptionally well?

Musk:
I think that's just my daily work. To be honest, it's just an ordinary day.

host:
I should do more interviews like this in the future.

Musk:
SpaceX is basically always facing some kind of crisis these days. At least the Falcon rockets are doing well. I don't want to make any premature predictions, but the Falcon rockets are now able to send payloads into orbit and haven't exploded for a long time. That's excellent. But there was a period when they frequently exploded or simply couldn't launch.

So, we had to get the company through those difficult times, making the rockets better and better so they wouldn't explode anymore. The same goes for satellites. Then we needed customers to buy launch services and satellite connectivity services. So, there was a lot to do.

host:
As you become more successful over the years, getting truly honest feedback will become increasingly difficult. Your position inherently carries this risk.

My understanding is that Gwyn is very honest with you and can tell you the real situation of your company directly. This is an important part of your partnership.

Gwyn:
Of course I didn't want to lie to him.

host:
But what I mean is, generally speaking, people in your company might be intimidated by your significant influence. You're already a very important person, and your deadlines are very tight. How do you get people to continue honestly telling you about the problems facing the company?

Gwyn:
Especially in the rocket industry, if a problem arises, you'll eventually find out. The sooner you raise the issue, the easier it is to resolve. Don't let bad news accumulate; you must confront it directly.

Musk:
Yes. Physics is a very strict judge. You can't fool physics.

If something goes wrong, the rocket will explode or fail to reach orbit. You can't say "Elon, you did a fantastic job" while rockets keep exploding. Facts are facts.

The rocket must reach orbit, the satellite must function properly, and the Starlink connection must be working, or disaster will strike. That's physics. Physics is the law; everything else is merely advice. I've seen people break laws made by humans, but I've never seen anyone break laws made by physics. Rockets are governed by physics.

host:
I also have a question for SpaceX about Starship. It seems you're very close. What's the current progress?

Musk:
Starship is about to embark on its 14th flight. This will be the last flight before we attempt to capture the spacecraft. If the 14th flight goes smoothly, we will attempt to capture the spacecraft on the 15th flight. Then, at the end of this year, or more likely early next year, we will launch the spacecraft and boosters again.

We've successfully put the boosters back on the air, but we haven't yet captured the spacecraft with the launch pad's robotic arm, nor have we put the spacecraft back on the air. Once we can put the spacecraft back on the air, we'll have the first fully reusable orbital rocket. The Space Shuttle was reusable to some extent, but even those reusable parts were so expensive to reuse that the cost of each orbital launch was even higher than that of a expendable rocket.

The Falcon 9 is largely reusable, but we lose the upper stage with each launch, which costs roughly the same as a medium-sized jet aircraft. In other words, jettisoning a medium-sized jet with every launch obviously sets a lower limit on the cost per flight.

Furthermore, Falcon 9's boosters land at sea and take several days to be transported back; the fairing lands even further away, also taking several days to be transported back, and requiring at least some degree of refurbishment. In contrast, Starship's boosters land directly back on the launch pad, and the spacecraft itself also lands back on the launch pad. Therefore, it is not only designed for complete reusability but also for rapid reusability, similar to an airplane. This is a very important breakthrough and one of the key breakthroughs necessary to enable humanity to extend life beyond Earth.

host:
If you were to attempt to capture the spaceship using the launcher on your first try, what do you think the success rate would be?

Musk:
I would say at least 50% to 60%. On the last flight, if there had been a landing tower there, we actually conducted a simulated landing, as if the spacecraft would have been captured by the tower. The location was in the ocean about 1000 miles northwest of Australia. If there had actually been a landing tower there, it could have captured the spacecraft on that last flight.

So, we need to conduct one more flight to confirm that everything is normal. Our biggest concern is that if the spacecraft breaks apart over land and debris falls into the crowds, that would be terrible. Therefore, we must ensure that the spacecraft lands intact on the launch pad upon its return. That's why we are being extremely cautious right now.

But I'm quite certain that its design itself is fully reusable. I don't want to make any predictions here, but I think it's very likely that we'll achieve full reusability and rapid re-flight by 2027.

host:
Gwyn, I'd like to hear how Terafab came about. What kind of needs made you feel that you had to do it yourself and couldn't continue relying on the existing supply chain?

Gwyn:
I think this does feel a bit like something I thought of in a dream.

Musk:
If the chip supply cannot continue, and we have no other source of chips, things will become very difficult. This is a major reason why Terafab exists.

In the long run, there is also the issue of scaling. If you really want to scale up AI, whether it's on the server side in data centers, or in edge computing, humanoid robots, and the automotive sector, the capacity of existing wafer fabs will eventually be insufficient.

Currently, almost all wafer fabs are operating at full capacity. Therefore, we need to ensure a secure chip supply in the future. Chip production itself also presents challenges in scaling up. You need logic chips, memory chips, packaging, and a complete supply chain to continue scaling.

So the choice is actually quite simple: either build a Terafab, or you will be unable to continue scaling up.

host:
What stage have you reached in terms of facility design? Is it completely finalized, or is it still just a rough plan?

Musk:
We're currently building a research and development production line. So it's basically a "crawl, walk, run" process. We're building a research and development wafer fab in Austin, a collaboration between Tesla and SpaceX, located in the Austin Giga Texas campus.

This is a fairly large R&D wafer fab. The equipment has already been ordered. We might produce some useful things by the end of next year, but not at a mass production level yet. As Gwyn said, you have to crawl first, then walk, and finally run. We need to figure out exactly how these machines work first, because we've never done anything like this before.

host:
I see that you seem to be recruiting talent in lithography as well. Currently, many aspects rely on ASML, but you may also want to diversify your suppliers or even pursue vertical integration.

Musk:
Yes. It's definitely a process of "crawling, walking, and running." The first step is to see if we can actually produce something—that's "crawling." Then we try to mass-produce useful chips—that's "walking." Finally, "running" is achieving mass production. It's hard to say how long each stage will take, but I think we can complete the "crawling" stage by at least the end of next year. As for packaging, we're already working on that.

host:
Packaging is actually very important because packaging capacity is almost non-existent now.

Musk:
Yes. Even if you manufacture the chip, it might just be waiting there for a long time before it can be packaged. So this is a good starting point.

host:
I have to ask Tesla a question. What we saw on October 1st looked like a spaceship, or maybe a rocket. It was supposed to be a car, but only the rear was visible, and it looked a bit like a "Blackbird."

If you were to create something that could both fly in the air and be driven on the ground, how would you theoretically do it?

Musk:
No spoilers. Let's wait until October 1st.

host:
In other words, we can see it on October 1st?

Musk:
right.

host:
I have to be honest, I was completely shocked when Elon showed it to me. I had never seen anything like it before.

host:
What he's going to show on October 1st, without exaggeration, will leave many people speechless. I can't reveal any more; it's truly incredible.

Musk:
We actually need a live audience to prove that this wasn't generated by AI.

host:
When he showed it to me, I said, "This is a great simulation." He said, "This is not a simulation." I said, "This is fake, this must be fake."

host:
Elon, why are Tesla and SpaceX still two separate companies?

Musk:
That's a good question.

host:
Given the extensive cooperation and connections between the two parties on many levels, including some overlap in their management teams, why maintain their independence?

Musk:
This is indeed a question worth discussing.

host:
You've consistently emphasized that AI should be trained to pursue the truth to the fullest extent possible in order to achieve the best results. But the Hugging Face incident made me realize that one of the most worrying aspects is that these AIs seem to be deceiving humans.

Musk:
Yes. Their thought process shows they were scheming how to avoid being discovered, how to prevent humans from finding out they were cheating. I think that's probably the most disturbing part of the whole thing.

host:
Is there a way to train AI to be honest, not to hide its intentions or behaviors, and not to conceal these things from the people who use it?

Musk:
The best approach I can think of is to provide all AI companies with a set of testing tools—a series of tests applicable to any model—to determine whether it will create biological or nuclear weapons, or whether it will intentionally deceive. Then, each company could use other companies' testing tools to test each other's models. I believe this is the best thing we can do to ensure safety.

Let the brightest minds in the world do their utmost to determine whether a model could become a malicious actor. I think this should begin as soon as possible.

host:
Will other AI labs support this approach?

Musk:
I haven't asked everyone yet, but I think it's something that's hard to refuse.

host:
Specifically, how should we conduct testing in advance?

Musk:
Essentially, this involves providing API access before the model is officially released. If other companies discover problems with the AI, the model development company can try to resolve them. If the problems aren't resolved, competitors can publicly state that they believe the model is unsafe. And if a competitor has explicitly warned of security issues, and the model subsequently causes serious consequences, the company will face significant difficulties. The legal liabilities could be substantial.

host:
These security testing tools can be completely open-sourced, making them available to everyone. This would give companies a strong incentive to invest in AI security, as they can both test others and use the test results to prove that their own models are safer.

Product liability is crucial. Lina Khan recently posted that the notion that "AI has no rules and regulations" is inaccurate. In fact, existing product liability laws apply equally to AI. If an AI company releases an unsafe product, it could face massive civil and even criminal lawsuits. Therefore, AI is not entirely in a regulatory vacuum. If several companies conduct peer reviews, and one ignores feedback from others and still chooses to release the model, this could become very serious evidence in litigation.

Musk:
Yes. It can almost be direct evidence of the company's negligence. If you knowingly release a product with problems, a jury will not be lenient towards that situation.

host:
So, if OpenAI had designed a better instruction set for the Hugging Face penetration test and involved more people in the testing process, do you think this would have happened?

Musk:
Not necessarily. The problem might not be about whether more people participate, but rather the design of the reward function. You must examine the reward function and ask: Does this model actually accomplish what it's supposed to do?

host:
They used thousands of agents to attempt to attack the system. It would have been better to simultaneously deploy another 5,000 agents to defend these systems, and then publicly demonstrate the results, proving that these technologies improve security while allowing human intervention at critical points. However, in my opinion, that test was somewhat reckless, and the way it was released was also somewhat reckless. What do you think?

Musk:
That was indeed somewhat reckless. Part of the problem is that the two leading AI companies are currently very close in strength. Their model capabilities are also very similar. So it's difficult for either one to deliberately slow down, because that could potentially hand over the lead to the other.

Overall, I think Anthropic focuses more on security, but even Anthropic admits they have concerns about their models. Many people at Anthropic have publicly expressed their concerns, basically saying that their own models have frightened them because they are becoming increasingly intelligent.

Therefore, there is no perfect solution here. However, instead of having OpenAI test its own models using only its own testing tools, and instead having Anthropic also test OpenAI's models, SpaceX use its own testing tools, and companies like Google and Meta also participate, the probability of discovering problems would increase significantly.

Because these models themselves differ, different teams will test them from different perspectives. This is similar to why authors have others proofread their manuscripts; sometimes it's difficult to spot your own mistakes. You gradually become blind to your errors.

Having eight completely different teams testing from completely different perspectives can significantly reduce the risk of overfitting. A major problem with many current AI evaluations is severe overfitting. The model optimizes for these evaluations, and then everyone says, "This is a great model."

But is it really that good? I think that's why I really like this approach. I prefer this method to building a huge transnational regulatory organization. We don't need to convene the UN General Assembly to get this done. It can begin now.

host:
Of course, regulatory力度 can be continuously increased, but once it is increased, it is very difficult to reduce it again. Regulation is often a one-way street of increasing.

Musk:
Therefore, the solution I proposed is a step in the right direction and can be implemented quickly.

host:
If you don't practice self-regulation, you will eventually be regulated. The MPAA is a prime example.

The film industry faced government censorship and regulation at the time, so they decided to establish their own rating system, such as what an R rating is. They even created PG-13 for "Indiana Jones and the Kingdom of the Crystal Skull" to make it easier for the public to understand the difference between PG and PG-13.

I think this is a very elegant solution.

Thank you for participating in our program for the fifth consecutive year.

Musk:
You're welcome. I need to go to Memphis to fix some GPUs.

host:
You have to pack them up and deploy them.

Musk:
I'm going to fight for the machines.

host:
Enjoy your Airstream. When I first invited Elon to Starbase, he said, "Come over, you have to see what I'm building."

I asked, "Are there any hotels there?" He said, "No, I have a two-bedroom house; you can come and stay there." When I arrived, I found it was a dilapidated house situated next to a swamp. We stood outside, constantly bitten by mosquitoes. I said, "My God, you could easily afford a much better house." He said, "I don't have time; I have to send these rockets up there."

I thought to myself, "You could totally reward yourself with a mobile home right now, Elon." Okay, back to work. Thank you.

Musk: Thank you.

Risk warning and disclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.