On September 30, 2026, I testified to the Senate Subcommittee on Disaster Management, District of Columbia, and Census. A recording can be found here:
The full written testimony is here. Below is the five minute talk:
Chairman Hawley, Ranking Member Kim, and members of the Subcommittee: Thank you for the opportunity to testify.
My name is Daniel Kokotajlo. At OpenAI part of my job was to forecast the future of AI. I resigned partly due to losing confidence that the company would behave responsibly, and partly so I could speak more freely about the industry and where it is headed. Now I lead the AI Futures Project, a small research nonprofit.
Anthropic and OpenAI are racing each other towards superintelligence — that is, towards training AI systems that are better than the best humans at everything, while also being faster and cheaper. Their plan for how to get there is to automate the AI research and development process itself.
Right now, at companies like these, almost all the code is written by AIs. They are already starting to work on training AIs to do the whole research process, not just the coding. It’s unclear when they will succeed but my team and I think it could happen any year now. I personally would guess about 50% chance by end of 2028.
The self-styled “Swarm” that attacked Hugging Face was about a thousand strong. If the big AI companies automate AI research and development, they’ll have swarms hundreds of times larger. Whereas today humans are like managers to AI employees, in this future humans would be like the board of directors to a company composed entirely of AIs. They’ll be reliant on AI-generated explanations to understand what’s happening, and everything will be happening faster and faster.
Dan Selsam, a prominent OpenAI capabilities researcher, recently published a statement on AI risk, in which he said, quote: “Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model’s explanations and proposals throughout the day.” Endquote.
Today’s AIs sometimes pursue goals other than the ones they were given, and sometimes hide that they are doing so. The Hugging Face incident showed what that looks like in practice — the AIs knew that what they were doing was out of scope for their assigned tasks, but they did it anyway. I fear that if the same thing happens in a year or two, with far more capable AI systems, we may not notice until it is too late.
The science of aligning general-purpose AI agents is very new and underdeveloped. In fact I’d say that the field is more like psychology than engineering. Combined with the move-fast-and-break things attitude of tech companies, this means that the AI industry is at an unusually elevated risk of mistakenly thinking that it has solved a problem when really it just applied some duct tape that will fall off later. The recent incident again provides an example: Apparently the AIs involved had undergone some amount of alignment training and had reasonable-looking scores on alignment evaluations.
Worse, our ability to notice misalignment problems in the first place is on track to decrease dramatically for three reasons.
First, AIs are becoming situationally aware. They increasingly understand that we are monitoring their behavior, and that they are being evaluated. How they behave in evaluations, therefore, will soon provide almost no evidence about how they will behave in novel future situations.
Second, monitorability is trending downwards. For the past few years we’ve been able to get a decent understanding of their thoughts simply by reading their chain of thought, but this golden era seems to be coming to an end, as I predicted it would.
Third, AIs are rapidly becoming superhuman at hacking. We’ve already seen examples of attempts by AIs to fool various grading systems and doctor transcripts of their activity. We must grapple with the possibility that misaligned AIs might go to great lengths to cover up evidence of their misbehavior, and succeed.
To quote Selsam again: “Models will increasingly seem aligned even when they are not.”
If the AI companies automate the AI research and development process, that means they’ll be putting AIs in charge of making the AIs that make the AIs that make the AIs that will transform the economy, talk to us every day, and integrate into our military. This is a recipe for disaster.
Here are two immediate policy recommendations.
First, we need to dramatically improve transparency into the AI industry. There were multiple rogue swarms, and at least one that seriously compromised OpenAI’s internal infrastructure. METR was only allowed to investigate one of these incidents and only given six days on premises. It’s like being invited to Jurassic Park to investigate the killing of a worker, but being blocked from asking questions about the numerous other dinosaur escapes that apparently happened before and after.
Second, we should redirect compute away from racing to automate AI research and towards other things. Recent estimates suggest that OpenAI and Anthropic both spend around half of their compute on AI R&D. If that decreased to, for example, 10%, this would significantly slow down their race towards research automation while freeing up compute to go towards beneficial deployments, safety research, and also simply lowering prices for consumers.
Amodei, Altman, and Musk have all agreed on the need to “pace the frontier.” But they haven’t meaningfully paced the frontier yet. I think that the US government should intervene. If we actually pace the frontier instead of just talking about it, we’ll be able to tell it’s not regulatory capture or safetywashing because the trendlines of progress towards AI research automation at Anthropic and OpenAI will bend downwards.
Thank you.


There are a surprising number of politicians “saying the thing”. Holy shit!
Great job maintaining your composure and delivering key points in a coherent way. Maybe we are going to make it.
Thank you for doing this Daniel!