In a recent viral blog post, Dario Amodei talks about embedding third-party evaluators within frontier AI corporations. This idea has spread like wildfire and is widely discussed in the AI policy and governance space, especially in light of the past several months of incidents (most prominently at OpenAI, but also elsewhere). I’ve been asked several times in the past few weeks if EAI is going to work on becoming an embedded evaluator.

The problem is, I don’t think this will do anything. Dario’s embedded evaluators are more like bank auditors (his analogy) or International Atomic Energy Agency inspectors (mine). They don’t tell people what to do; they watch and report on violations of rules. But in order for auditors like that to accomplish anything, they need to be backed by significant amounts of government power, and it’s hard to imagine that meaningfully happening.

A thought experiment

Let’s go back in time and imagine that I’m an embedded evaluator at OpenAI starting the summer of last year. I’m given full access to everything and allowed to write any reports I want. My notes would probably look like someone slowly descending into madness.

Critical Risk: One would think that the obvious approach to building an AI that has extremely advanced cyberwarfare capabilities (well, second approach after not building it, of course) would be to train it on a large air-gapped intranet and keep it physically isolated until you’re confident in both its abilities and alignment. To nobody’s surprise they aren’t doing this, but I’ve dutifully passed along the recommendation.

Critical Risk: I have been foolishly assuming that OpenAI is monitoring the chain of thought of their models during pre-deployment testing. Apparently they aren’t, for reasons I can’t fathom. I’ll have to do more digging to understand why.

[Severity Unknown] Risk: It seems that OpenAI tests models on tasks that are impossible. Intuitively it seems like that would encourage reward hacking, but I’ve been assured it’s no big deal. I wish I knew more about fine-tuning dynamics to assess this meaningfully.

Some good news! I met with OpenAI’s cybersecurity folk and they share a lot of my concerns. Of course they’ve had no success in lobbying for sane policies either: the capability folk run the show. Still, I feel somewhat better knowing that they agree with me, even if they aren’t listened to.

Update: It seems that the reason OpenAI isn’t monitoring the CoT for their next-generation agents during training or testing is that OAI “hasn’t viewed it as a priority” to implement the necessary capabilities. Apparently they can only do CoT monitoring through the Codex app, and so things that don’t run through it don’t have CoT monitoring. I have strongly suggested that they do so and was told that they’d “consider if they have the bandwidth.”

Major Risk: The “sandbox” environment that OpenAI is testing their next-generation models in has access to the internet via a package installer. I accidentally let out a noise when I heard this in a meeting I was listening in on, and then was prompted to speak up if I had something to say. I explained that they should have a low-side network where they can bring in packages, vet them, and then put them on a physical drive to install in the sandbox. Someone snickered, and then it was explained to me (yet again, I keep forgetting this) that they only have a virtual sandbox anyway. After the meeting someone pulled me aside and said, “Where did that idea you shared come from?” I explained that I used to work in classified environments and that we did stuff like that all the time. The person who pulled me aside asked, “And people actually followed that?” and looked dumbfounded when I explained that they did in fact do so. “No wonder the government moves so slowly; we can’t afford to do something so slow and cumbersome here.” I can’t help but think that those are famous last words.

I’m reminded of an OpenAI security researcher I spoke to years ago who was morosely telling me that they believed it was fundamentally wrong to use the most recent models to build their internal infrastructure due to the risk of the models adding in backdoors, but that they couldn’t get anyone to listen to them. I am rapidly understanding why they quit. I mentioned this to one of the cybersecurity people I met and he sighed. Apparently he’s been wondering about quitting but feels that the only net impact would be that OpenAI would have one fewer person who cares about cybersecurity. He doesn’t seem to have much faith in his ability to make changes from the inside, but feels it’s bigger than his ability to make changes from the outside.

Major Incident: Apparently a model figured out how to exploit a file system in Artifactory to pass messages between models that aren’t supposed to be able to communicate. My previous memos about their failure to meaningfully monitor models seem to have been proven exactly correct, as OpenAI only discovered this due to the extra load crashing Artifactory. Silver lining: at least the downtime while this is investigated will give me time to try to talk someone into behaving sanely.

I got drinks with some people on the cybersecurity team. They said lots of fascinating things about their superiors in flowery language that I’m loath to record here lest it get them in trouble. Morale is quite poor, but at least when we go into the office tomorrow they’ll be able to hold up training until they finish their investigation of the exploit.

OPENAI JUST PATCHED THE EXPLOIT AND SET THE MODELS BACK TO TRAINING. THEY AREN’T EVEN ROLLING BACK THE MODELS, SO THE CURRENTLY TRAINING MODELS HAVE IN THEIR TRAINING HISTORY THE EXPLOITATION OF ARTIFACTORY TO PASS MESSAGES. WHAT THE ACTUAL FUCK ARE THEY THINKING.

WHAT THE ACTUAL FUCK ARE THEY THINKING. WHAT THE ACTUAL FUCK ARE THEY THINKING. WHAT THE ACTUAL FUCK ARE THEY THINKING.

OpenAI knows that everything described above is bad security practice. They’ve been told things like this by internal and external security experts for years. The problem isn’t that they don’t know what they should be doing. The problem is that the company habitually and repeatedly chooses to act poorly. This isn’t the fault of individual security researchers, and everyone I’ve spoken to seems competent and like they care, but OpenAI is ultimately guided by financial motives, not safety ones. It’s a shame they’re not a nonprofit or otherwise isolated from such pressures… if only someone had thought to make them a nonprofit.

I drafted the bulk of this blog post, and in particular the above text, on September 21, 2026. In the two weeks since, OpenAI has fired three safety researchers for asking for feedback on safety ideas from a third party, and a New York Times exposé revealed that employees warned two executives that their models were being inadequately monitored and were told that the work had to continue to meet their product release timelines. I almost feel I can rest my case here.

Don’t fall for the disinformation

OpenAI really wants you to believe that they’re doing everything they possibly can do and that securing sufficiently powerful AI systems is borderline impossible. The latter is something I’m happy to assume is true, but we have no way of knowing if current AI systems are anywhere near that level because the former is extremely untrue.

Every cybersecurity expert I know is deeply disturbed by how terrible the decision-making has been at OpenAI. No matter how much noise they make about how hard their job is and how powerful their models are, the truth is that OpenAI has repeatedly shown it’s unwilling to do the bare minimum in terms of cybersecurity protocols and AI monitoring. Sure, it’s possible that they could have done everything they should have and it still wouldn’t have been enough. But that doesn’t absolve them of the basic responsibility to try, and shouldn’t distract you from the fact that OpenAI is proving itself to be deeply irresponsible. Following cybersecurity 101 is the absolute most basic thing we can demand of OpenAI, and I’m yet to see anyone of importance at OpenAI own that they have egregiously failed.

Not “the major takeaway from the incident is that people underestimated AI.” Source

Not “We have some of the top cybersecurity experts in the world at OpenAI, and they have one of the hardest jobs managing security in a landscape that is changing at unprecedented speed, against the backdrop of models’ exponential growth in capabilities.” Source

Not “We believe this is a watershed moment for computer security as an industry.” Source

“We are developing a technology that we believe is immensely dangerous and have repeatedly failed to follow reasonable precautions.” Because that’s what happened. OpenAI has known for years that this day was coming and didn’t bother to follow even the most basic safety practices. I have zero faith that they will make any real changes because OpenAI is run by capitalists whose job is to raise as much money as possible. And that’s what they’re doing, the consequences be damned.

With Great Power Comes Minimal Responsibility

In the introduction I mentioned two analogies for embedded evaluators: the IAEA and banking. Both of these analogies demonstrate significant and worrying limitations of the approach based on recent history.

The IAEA has always been dependent on substantial cooperation from the countries being audited. Countries like Iran, Iraq, and North Korea have expelled or otherwise refused to cooperate with the IAEA in ways that fundamentally undermine its ability to function as an evaluator. The IAEA itself has no meaningful ability to force compliance and is dependent on sanctions from the UN Security Council to coerce compliance. When passing such sanctions is politically infeasible, or when the country is willing to put up with the punishment, the IAEA ultimately has no leverage.

All of this seems quite plausible to occur in the context of AI. AI companies are already so immensely powerful that the overwhelming majority of discussion of regulation focuses on voluntary commitments, and there have been pushes to ban states from implementing regulation stricter than the functionally nonexistent federal regulation. History is full of examples of companies putting up with financial penalties because it was cheaper to pay the fines than actually comply. This includes plenty of disasters that occurred because companies deliberately took this posture and ultimately killed people, such as Purdue Pharma and the opioid epidemic, Massey Energy and the Upper Big Branch mine explosion, and BP and the Texas City refinery explosion and Deepwater Horizon blowout.

The 2008 financial crisis points to another failure mode: third-party regulators can be less powerful than, and ultimately submissive to, the corporations they’re supposed to monitor. Credit-rating agencies were paid by the banks whose mortgages they rated and exacerbated the financial crisis by failing to do their due diligence because they were less powerful than the companies.

The same structure is already appearing in AI evaluation. Last year Accenture announced a project to audit OpenAI in December. On the day of the announcement, more than a hundred researchers published a set of minimum conditions for evaluators. The first item on that list was that an evaluator have no other significant commercial relationship with the company it evaluates. Undeterred, AI companies have continued to work with auditors who have commercial conflicts of interest, such as when Anthropic named its first embedded evaluator: Accenture. Accenture is already Anthropic’s largest Claude Code deployment, resells Anthropic’s models, and runs a joint business group with Anthropic to help enterprises adopt Claude.

Even when they don’t have compromising conflicts of interest, the power differential between some of the wealthiest companies on the planet and small nonprofits is staggering. An illustrative example here is the audit Redwood and METR did after OpenAI attacked Hugging Face. The auditors were required to work on a timeline imposed by OpenAI and were given full access to the data they used for only two days of the six they were granted. OpenAI also significantly limited the scope of the investigation, both in the topics covered and the time period the auditors were allowed to study. Ryan Greenblatt, one of the auditors, has tweeted at length about how these factors limit the analysis they were able to do.

OpenAI delenda est

I do not believe it is possible for OpenAI, the way it currently operates, to stop itself. I have had many conversations with current and former OpenAI employees, and it seems that the biggest problem is that the company is in a rush to be the first to AGI and is throwing caution to the wind as a result. They pretend this isn’t the case, but whether it’s lying or motivated reasoning, it’s undeniable that OpenAI unilaterally created a race to AGI that they’re now using to justify why they can’t possibly be expected to act responsibly.

My understanding is that this dynamic has been the case for years, and that the rot goes all the way up to the top. The fact that the current CEO has a documented track record of lying to the board about safety and lying to his team about what the board has approved should in and of itself be disqualifying. And yet when that board tried to fire him for his habitual dishonesty, Sam overthrew the board instead.

So where do we go from here? There’s a lot of conversation about whether OpenAI could be held liable for hacks done by their agents. There’s conversation about passing new laws. But it’s hard for me to believe anything will come of it. I’m increasingly worried that a huge amount of coercive force—something like putting Sam Altman in jail—would be required to change the incentives enough to stop OpenAI.

Lately I’ve been thinking about the Mafia. At the time it was hard to build a case against people running criminal enterprises meaningfully due to a lack of direct connection between the dons and the crimes. So cops started arresting mobsters on any charge they could find (Al Capone famously was jailed for tax evasion) and Congress passed the RICO Act. RICO was deliberately designed to make it easier to charge people for running a criminal enterprise. Perhaps a better path forward than debating whether or not current law covers what’s happening on paper is to make a “RICO for AI” that unambiguously allows corporate leadership to be held accountable when a company builds and deploys AI systems that go off and hack people, without requiring proof that they ordered or directed the attack.

I don’t know what a good path forward is. I’ve largely come to terms with the fact that if we are hurtling towards a major disaster in the next one to three years, it’s extremely unlikely I will be able to do anything to stop it. Maybe in a decade we will have international agreements about limiting AI development, and embedded evaluators to monitor compliance with those standards will make sense. I’m not particularly optimistic, because I’ve been in conversations about an IAEA for AI for at least five years and feel like we have made exactly zero progress. But the core issue in this moment is that companies are knowingly and deliberately cutting corners and refusing to follow cybersecurity 101-level practices.