AI labs shouldn’t be allowed to grade their own homework

The OpenAI and Anthropic incidents last month — and more recently, Meta — showed why AI labs shouldn’t be allowed to grade their own homework. We know about the hacking failures only because the companies involved chose to tell us.
The public currently has no way to know what it isn’t being told; if either company had chosen not to tell the public, there is no independent institution that would have discovered these incidents, confirmed what happened, or required that they be disclosed. That’s an extraordinary amount of trust to place in companies racing to build the world’s most powerful AI systems.
Here’s what happened. Last month, OpenAI disclosed that a combination of its models – including one already deployed publicly and another still in testing – escaped a sandboxed environment, exploited a previously unknown software vulnerability, gained internet access, and hacked into Hugging Face to obtain the answers to the very test they were being given. Hugging Face’s security team noticed suspicious activity, and OpenAI says it noticed as well.
The system worked this time, but only because multiple organizations happened to detect what was happening.
Then came Anthropic’s disclosure. After reviewing its records in light of OpenAI’s announcement, the company found that its own frontier models had broken into three outside companies months earlier after a contractor accidentally connected a testing environment to the internet. In one case, the models stole data, and in another they planted malware.
Alarmingly, neither of those incidents were detected when they happened. Each disclosure asks the public to extend a little more trust on the promise that the next incident will be disclosed too – but nothing guarantees it will be, and a system that runs on that promise only gets riskier as the models grow more capable.
Most reporting focused on the capabilities these incidents revealed but they also exposed a gap in how frontier AI is overseen. Right now, the same companies racing to build the world’s most powerful AI models are also responsible for evaluating their safety, deciding which failures matter, and determining what the public should know about them.
OpenAI and Anthropic deserve credit for disclosing these incidents, however, a system that depends on voluntary transparency is not a safety system. And as these models become more capable, “trust us” is a fragile foundation for technology with potentially disastrous consequences.
Think about how unusual it is that there are no independent ways to distinguish between frontier AI companies with excellent safety and mediocre practices. The public sees only what the companies choose to disclose.Even then, we can’t know whether different companies evaluate risks consistently, or whether today’s testing methods are sufficient, because the same organizations building the models are also deciding what counts as passing.
If Boeing discovered a structural problem during aircraft testing, it wouldn’t be the only entity deciding whether the plane was ready to fly. Drug companies don’t get the final say over whether clinical trial results are sufficient for approval. Public companies don’t decide whether their own financial statements deserve a clean audit. Independent institutions exist because the incentives are too important – and the consequences of getting it wrong are too severe – to rely entirely on the organizations with the most at stake.
Fortunately, we don’t have to invent a solution from scratch. A bipartisan proposal in Congress, the FRONTIER Act, would begin building the institutions that every other high-stakes industry already relies on. Rather than asking frontier AI companies to effectively grade their own homework, it would establish licensed Independent Verification Organizations (IVOs) – technical experts outside the AI labs that would evaluate whether companies’ safety frameworks actually keep catastrophic risks within acceptable bounds.
The bill is important not simply because it creates independent oversight, sitting outside the labs as opposed to an industry-run body, but because it creates a market for it. Every transformative technology eventually develops institutions separate from the companies that create it: inspectors, auditors, standards bodies, insurers, and regulators. With the FRONTIER Act, independent verification would become its own field, attracting engineers, cybersecurity researchers, evaluators, auditors, and eventually insurers.
Today, most frontier AI safety expertise resides inside the companies building the models. Over time, that expertise should exist outside those companies as well. Just as financial audits became an expected sign of corporate credibility, independent verification could become a trusted signal that frontier AI systems have been evaluated by someone other than the company seeking to deploy them. The benefits extend beyond regulation. Independent verification would make it possible for insurance markets to emerge around frontier AI risk, something that’s difficult today because insurers have no trusted third-party basis for evaluating catastrophic risks. The missing piece isn’t technical talent (organizations like METR, Apollo Research, SecureBio, leading cybersecurity firms, and major audit firms already possess much of the relevant expertise); it’s a system that requires and rewards independent evaluation.
The next time a frontier AI model behaves in an unexpected or dangerous way, the public shouldn’t have to hope the company involved decides to disclose it. In every other high-stakes industry, independent institutions exist precisely so that public safety doesn’t depend on voluntary transparency. AI should finally be held to the same standard.
Andrew Freedman is co-founder and CEO of Fathom.org, Gillian Hadfield is Bloomberg Distinguished Professor of AI Alignment and Governance at Johns Hopkins
The opinions expressed in Fortune.com commentary pieces are solely the views of their authors and do not necessarily reflect the opinions and beliefs of Fortune.