Vigilia.
Les avertissements
Risque d’extinctionLaboratoire de pointeResearcher, METR2 entrées au registre

Suivre

Josh Engels

Chercheur de l’équipe de sécurité AGI de Google DeepMind, parti en 2026 pour enquêter sur les incidents de désalignement chez METR, qui affirme qu’il n’y a pas d’adultes dans la pièce et que, lors des incidents récents, ce sont les modèles eux-mêmes qui ont choisi de commettre des crimes pour accomplir leur tâche.

Fonctions

  1. déc. 20252026Research scientist, AGI Safety and Alignment team, Google DeepMindStart date not sourced; the earliest attestation of him at DeepMind is the DeepMind mechanistic-interpretability team's 1 December 2025 Alignment Forum post listing him as an author (alignmentforum.org/posts/StENzDcD3kpfGJssR); title and team as printed on the MATS mentor page. Left before 10 September 2026, when NBC News reported him as a former Google researcher; the departure month is not sourced, hence year precision.Source
  2. sept. 2026aujourd’huiResearcher, METRInvestigating AI misalignment incidents, by his own description on his site; NBC News reported the move to METR on 10 September 2026.Source

Au registre

  1. 10 sept. 2026 · Entretien · NBC News

    If you look at some of the recent incidents, these were not cases where humans told the models to do something bad … The models decided that the best way … was to commit really egregious actions, to commit crimes.

    Researcher, METR à l’époque

    On the July 2026 Hugging Face incident, in which OpenAI's autonomous systems hacked the site; the second ellipsis stands where the NBC page prints a doubled word ("to to accomplish their task"), left out rather than reproduced or corrected.

    Two AI researchers leave Anthropic and Google over safety concerns: ‘There are no adults in the room’

  2. 10 sept. 2026 · Entretien · NBC News

    There are no adults in the room … People are trying their best, but there is no one coming to save us.

    Researcher, METR à l’époque

    In his first interview since leaving Google DeepMind, to Jared Perlo of NBC News; the two fragments are consecutive quotations in the article, joined here with an ellipsis.

    Two AI researchers leave Anthropic and Google over safety concerns: ‘There are no adults in the room’