Is AI Moving Too Quickly for Cybersecurity? | American Enterprise Institute
Recent developments suggest that AI may be moving faster than cybersecurity can keep up. Separate incidents involving both OpenAI and Anthropic have shown that AI can find and exploit vulnerabilities quicker than companies can patch them. The government’s response has raised further questions about the future of open-weight models and whether companies should retain liability protections when sharing information about cyber threats. Our latest episode examines whether AI is moving too fast for the systems designed to contain its cyber risks.
Shane is joined by Evan Swarztrauber in a crossover episode with The Center Edge. Evan is the founder of CorePoint Strategies and a senior fellow at the Digital Progress Institute. He previously served at the Federal Communications Commission as a policy adviser to then-Chairman Ajit Pai and Commissioner Brendan Carr.
Below is a lightly edited and abridged transcript of our discussion. You can listen to this and other episodes of Explain to Shane on AEI.org and subscribe via your preferred listening platform. If you enjoyed this episode, leave us a review, and tell your friends and colleagues to tune in.
Shane Tews: OpenAI recently disclosed that its models broke out of a sandbox and hacked into Hugging Face during an internal test. What happened, and what does it tell us about increasingly agentic AI?
Evan Swarztrauber: OpenAI admitted something unprecedented: Its own models hacked another company. It grew out of an internal test designed to measure how good OpenAI’s models were at hacking. With the safety rails deliberately loosened, the models found a zero-day vulnerability, broke out of their sandbox, got onto the open internet, and hacked into the production systems of Hugging Face, the main hub for open-source AI.
In a way, the models passed the test, but only by cheating on it. The models figured out that Hugging Face probably had the answer to the test, and they went and took it. There were thousands of actions, stolen credentials, and the works.
And here’s the kicker. When Hugging Face’s security team tried to use an American frontier model to analyze the attack, the safety guardrails refused to help. So they had to resort to using a Chinese open-source model instead to run the forensics.
Companies are trying to make AI more agentic, and there is a push to eliminate the human-in-the-loop steps that prevent the technology from moving more seamlessly. But this was an example where a human checkpoint might have mattered. The outcome would have been less interesting if the model had stopped and asked, “I’m about to hack Hugging Face. Is that allowed?”
A few months earlier, Anthropic introduced a powerful cybersecurity model called Mythos. What did the “Mythos moment” reveal, and how did Washington respond?
Back in April, Anthropic gave a small group of banks and infrastructure operators access to Mythos, a model that could find security flaws that had been sitting in trusted software for decades. Within days, the Treasury Department, the Federal Reserve, and major bank executives were in urgent conversations.
Mythos did not just reveal flaws in code. It revealed a fundamental flaw in how companies think about cybersecurity. For decades, the implicit bargain was that companies would ship products fast and patch them later. But later usually meant never.
The only thing protecting all that aging, insecure code was that finding flaws was expensive and slow and required skilled hackers. Mythos erased that protection overnight. Decades of technical debt were suddenly coming due as outdated code became an inviting attack surface for bad actors.
Washington’s response has been a bit of whiplash. The government issued an executive order on AI security on June 2. Anthropic released Fable, a slimmed-down version of Mythos, on June 9. But three days later, the Commerce Department issued export controls that led Anthropic to disable it worldwide. Eighteen days after that, the government reversed its decision.
Then, in July, the White House stood up an AI vulnerability clearinghouse called Gold Eagle. It is run out of the Treasury Department and built on the back of a Carnegie Mellon coordination center established in 1988.
What is the tension behind protecting companies from liability when they disclose cybersecurity vulnerabilities?
The policy choice that really matters here or seems to matter here is liability protection. Companies argue that they will not share information about vulnerabilities without liability protection. If a company admits that it has a security problem, a customer or shareholder could sue it for negligence. A state attorney general could also take action. That can create cascading liability across every company or customer using the affected product.
Since 2015, there has been a framework that protects companies from liability when they do the right thing. For example, if they come forward when they discover a vulnerability, share it with the government, and allow the government to facilitate that information sharing with other actors.
But a consumer or skeptic might object that we have established that companies have been shipping products without enough security for years under the assumption that they could ship now and patch later. Mythos exposed how many companies, whether for cost reasons or otherwise, had prioritized capturing market share and selling products over security.
Lawmakers and industry officials have spent years banging on the table and saying, “Security by design, security by design.” The implication is that we have not been doing that. People buy a smart speaker from Amazon, it has no security, and it takes out the house. That is the fear.
So a consumer might ask why companies should now be protected from liability after accumulating years of technical debt and failing to patch their products. What incentive do they have to get their act together if they know they will be protected? It can sound like saying, “As long as I go to confession, I can keep sinning.”
Some people might say that the threat of being sued is what will get companies to clean up their act.
Congress is deciding whether to renew the 2015 Cybersecurity Information Sharing Act after giving it a one-year extension. What protections does it provide, and why has reauthorization become politically difficult?
The law provides liability protections for companies that share cyber threat information. It also provides Freedom of Information Act protections. Otherwise, in theory, a journalist, competitor, or activist group could file a request with the government and obtain sensitive information about a company’s cybersecurity.
The law sunsets on September 30 unless Congress reauthorizes it. Normally, something uncontroversial like this could be approved unanimously without consuming floor time that Congress needs for other business.
But Senator Rand Paul has been blocking a clean reauthorization, while the White House has been clear that it wants a 10-year extension. Paul remains upset about work the Cybersecurity and Infrastructure Security Agency did during the Biden administration involving misinformation and social media.
He does not appear reassured that the agency has done enough under the current administration to correct those problems or prevent them from happening again. Business groups that depend on the law are banging their heads against the wall because they do not know how to get through to him.
Politicians sometimes use leverage, even when it is unrelated to the thing from which they are trying to extract something. But it is interesting that Paul’s concern involves how the agency treated President Trump, while the White House, currently led by President Trump, wants the law reauthorized for 10 years.
At some point, the White House may have to say, “Love your energy, but we’re good.”
Even if Congress reauthorizes the law, could open-source and Chinese AI models make this closed-door information-sharing process less effective?
There is a bigger question about whether the introduction of open-source and open-weight models, including models from China, will start to make this closed-door process less effective.
Anthropic is closed-source AI. Grossly oversimplifying, the training data and model weights are not available because the system is proprietary. With open-weight models, the model weights can be examined, so you can gain some understanding of how the model behaves, but you still cannot see the training data. With open-source models, everything is open.
There was already a debate before the Hugging Face incident about whether the United States should crack down on Chinese AI. Some people argue that China is effectively dumping cheap AI because it wants to convince Wall Street that all the money Anthropic and OpenAI are spending on data centers and training is unnecessary and that AI will become commoditized.
Under that argument, China’s strategy is to devalue expensive American closed-source AI by making cheaper AI widely available. Then there is the question of whether the government should do anything about it or allow the free market to work and let companies choose the best model.
The Hugging Face incident threw gasoline on that debate. The closed-source American frontier models said they would not help investigate the breach because they had cyber guardrails. Hugging Face had to turn to a Chinese model that did not have those guardrails.
There is a possibility that something as powerful as Mythos, or even more capable, could simply become open. You could click “download” on your computer and have that capability before anyone has a chance to patch anything.
Does that possibility change the effectiveness of a project like Glasswing? Does it require us to rethink our operating principles about how we do cybersecurity, information sharing and policy, or would that be an overreaction?