Vigilia.
The warnings
Extinction riskFrontier lab6 entries on the record

Follow

Joe Benton

Manager of Anthropic’s Scalable Oversight team who left in August 2026 for the independent evaluator METR and wrote that companies racing toward self-improving AI could make progress uncontrollable, that humanity may not survive the transition, and that a lab could lose control of its systems without the public ever knowing.

Roles

  1. Nov 2023Jul 2025Member of Technical Staff, AnthropicModel organisms of misalignment, chain-of-thought monitoring and control evaluations for the Responsible Scaling Policy, by his own CV.Source
  2. Jul 2025Aug 2026Manager, Scalable Oversight team, AnthropicAlso research lead for the Anthropic Fellows Program; the end date is at month precision from his own essay of 11 September 2026 ("Two weeks ago, I left the safety team at Anthropic"), and he announced he would soon join METR as an independent evaluator, a role not yet started and so not listed.Source

On the record

  1. 11 Sept 2026 · Post · X

    AI companies are racing to build machines that are much smarter than any human, and we may not survive this.

    The post announcing his departure, at x.com/JoeJBenton/status/2098480585119572317, as carried in full by ANI via LatestLY; Parameter, Blockonomi and CoinCentral reproduce the same sentence, and Outlook India prints it with a closing comma.

    Joe Benton Resigns from Anthropic Safety Team, Researcher Warns AI Race to Superintelligence Risks Humanity

  2. 11 Sept 2026 · Post · Substack

    An AI company could undergo an intelligence explosion, or lose control of their systems, without the public ever knowing. I don’t think that is acceptable.

    From the passage explaining why he is joining METR; he writes he is pessimistic about political will for capability restraint while the frontier is invisible to everyone outside the companies.

    Why I left Anthropic’s safety team to hold AI companies accountable

  3. 11 Sept 2026 · Post · Substack

    If the pace of progress continues and the industry does not prioritize safety more heavily, I expect much worse to come: humanity could be permanently disempowered by the AI systems these companies build in the next few years.

    Written after citing the OpenAI agents’ Hugging Face incident, whose public disclosure he calls partly a matter of luck.

    Why I left Anthropic’s safety team to hold AI companies accountable

  4. 11 Sept 2026 · Post · Substack

    Within the next couple of years, we may be sharing the world with AI agents smarter than any human alive today. These AI systems may have drives and desires that diverge from those of any human overseer, with capabilities we can’t effectively constrain. Humanity may not survive this transition.

    The paragraph continues: “We need a lot more preparation to make this world safe.”

    Why I left Anthropic’s safety team to hold AI companies accountable

  5. 11 Sept 2026 · Post · Substack

    Frontier AI companies are racing to build AI systems that can recursively self-improve. The aim is to build “superintelligence”, or an AI system much smarter than any human. If they succeed at this goal, the rate of AI progress may go from merely fast to uncontrollable.

    His own essay, published the day he announced the departure; it opens by saying AI companies are on track to impose an unprecedented level of risk on society.

    Why I left Anthropic’s safety team to hold AI companies accountable

  6. 10 Sept 2026 · Interview · NBC News

    All of these companies — and this is something I witnessed firsthand at Anthropic — are pretty directly trying to race towards automating the process of AI R&D itself

    His first interview since leaving, with Tom Llamas, published 10 September 2026; he added that he is worried things might progress too fast for us to get our act together in time, unless we worry about it now.

    Two AI researchers leave Anthropic and Google over safety concerns: ‘There are no adults in the room’