About

What I work on: making sense of current AI safety work, the cruxes that decide what labs and funders do next, and the alignment of superintelligent AI.

I’m Jacques Thibodeau. I work on reducing risks from advanced AI, and more broadly on whether the shift to powerful AI goes well for people.

Most of my time now goes to three things. The first is finding the disagreements that decide what labs, funders and governments do next, and working out what evidence would settle them. The second is making sense of current AI safety work, which is tangled enough that people often argue past each other. The third is the alignment of superintelligent AI, including what models actually know rather than what they only appear to know.

The work

The alignment research dataset and paper. In 2022 I collected and cataloged the AI alignment research literature and analyzed it to find the field’s real subfields, rather than the ones people assert exist. We released a paper, “Researching Alignment Research: Unsupervised Analysis”, and an open dataset, and showed that a classifier trained on the corpus finds relevant work nobody had put in it. The project started at AI Safety Camp. Jan Leike, then at OpenAI, called this direction his “favored approach to solving the alignment problem” and mentioned the work in a post at the time.

SERI MATS. In July 2022 I went to Berkeley for two months on the SERI MATS program, to work on aligning language models and on making language models useful for accelerating alignment research.

The ROME result. ROME was one of the most influential model editing papers in prosaic alignment. I tested it and found the edit does not generalize the way people assumed. It is not bidirectional, it mostly edits the token association rather than the concept, and it over or under optimizes depending on the new fact. Read it. The wider point is that these interventions usually show correlation, not causation.

Automating alignment research. This was my main focus for a couple of years. It is now one of the areas I know well rather than the thing I work on. Nearly every AI safety plan includes “automate alignment research”, and almost nobody says which of four different things they mean. I wrote Gaining clarity on automated alignment research to separate them, because the disagreements people think are empirical are usually about which sense they have in mind. Automating AI safety: what we can do today lists concrete projects that would make current coding agents better at running safety experiments. That post came out of a SPAR project, and PIBBSS supported me while I wrote it. I also gave a talk on it at the PIBBSS Symposium 2025, which runs about an hour; that page has the video, the deck and the transcript.

Now. I’m starting a non-profit research organisation. It works on the disagreements that actually decide what labs, funders and governments do next, and on the evidence that would settle them or at least sharpen them. The output is reports, research and scenario planning, written for the people making those calls. It’s early, and quiet for now. If you work on this, get in touch on X or LessWrong.

How I think about this

I want a world with less unintentional suffering and no existential catastrophes. I work on AI because I think superintelligent AI is the most consequential thing humans will build, and it can produce either outcome. Above that, I want people to be able to meet their needs and to have room for the kind of experience Scott Barry Kaufman calls transcendence. That is the whole reason the technical work matters to me.

Disclosures

  • I received a grant from the Long-Term Future Fund in 2022 to continue the work on accelerating alignment research.
  • PIBBSS supported me during the writing of “Automating AI safety: what we can do today”.
  • In 2025 and 2026 I worked on an AI safety startup and decided not to keep pursuing it. Why.