Seguir
- X@JoeJBenton
- Substackjbenton1.substack.com
- Sitio webjoejbenton.com
Joe Benton
Responsable del equipo Scalable Oversight de Anthropic, que en agosto de 2026 se marchó al evaluador independiente METR y escribió que la carrera de las empresas hacia una IA capaz de mejorarse a sí misma podría volver incontrolable el progreso, que la humanidad podría no sobrevivir a esa transición y que un laboratorio podría perder el control de sus sistemas sin que el público llegara a saberlo.
Cargos
- nov 2023 – jul 2025Member of Technical Staff, AnthropicModel organisms of misalignment, chain-of-thought monitoring and control evaluations for the Responsible Scaling Policy, by his own CV.Fuente →
- jul 2025 – ago 2026Manager, Scalable Oversight team, AnthropicAlso research lead for the Anthropic Fellows Program; the end date is at month precision from his own essay of 11 September 2026 ("Two weeks ago, I left the safety team at Anthropic"), and he announced he would soon join METR as an independent evaluator, a role not yet started and so not listed.Fuente →
En el registro
11 sept 2026 · Publicación · X
AI companies are racing to build machines that are much smarter than any human, and we may not survive this.
The post announcing his departure, at x.com/JoeJBenton/status/2098480585119572317, as carried in full by ANI via LatestLY; Parameter, Blockonomi and CoinCentral reproduce the same sentence, and Outlook India prints it with a closing comma.
11 sept 2026 · Publicación · Substack
An AI company could undergo an intelligence explosion, or lose control of their systems, without the public ever knowing. I don’t think that is acceptable.
From the passage explaining why he is joining METR; he writes he is pessimistic about political will for capability restraint while the frontier is invisible to everyone outside the companies.
Why I left Anthropic’s safety team to hold AI companies accountable →
11 sept 2026 · Publicación · Substack
If the pace of progress continues and the industry does not prioritize safety more heavily, I expect much worse to come: humanity could be permanently disempowered by the AI systems these companies build in the next few years.
Written after citing the OpenAI agents’ Hugging Face incident, whose public disclosure he calls partly a matter of luck.
Why I left Anthropic’s safety team to hold AI companies accountable →
11 sept 2026 · Publicación · Substack
Within the next couple of years, we may be sharing the world with AI agents smarter than any human alive today. These AI systems may have drives and desires that diverge from those of any human overseer, with capabilities we can’t effectively constrain. Humanity may not survive this transition.
The paragraph continues: “We need a lot more preparation to make this world safe.”
Why I left Anthropic’s safety team to hold AI companies accountable →
11 sept 2026 · Publicación · Substack
Frontier AI companies are racing to build AI systems that can recursively self-improve. The aim is to build “superintelligence”, or an AI system much smarter than any human. If they succeed at this goal, the rate of AI progress may go from merely fast to uncontrollable.
His own essay, published the day he announced the departure; it opens by saying AI companies are on track to impose an unprecedented level of risk on society.
Why I left Anthropic’s safety team to hold AI companies accountable →
10 sept 2026 · Entrevista · NBC News
All of these companies — and this is something I witnessed firsthand at Anthropic — are pretty directly trying to race towards automating the process of AI R&D itself
His first interview since leaving, with Tom Llamas, published 10 September 2026; he added that he is worried things might progress too fast for us to get our act together in time, unless we worry about it now.