Vigilia.
Die Warnungen
AuslöschungsrisikoFrontier-LaborResearcher, METR2 Einträge im Verzeichnis

Folgen

Josh Engels

Forscher im AGI-Sicherheitsteam von Google DeepMind, der 2026 ausschied, um bei METR Fehlausrichtungs-Vorfälle zu untersuchen, und der sagt, es seien keine Erwachsenen im Raum und bei den jüngsten Vorfällen hätten die Modelle selbst entschieden, Straftaten zu begehen, um ihre Aufgabe zu erfüllen.

Rollen

  1. Dez. 20252026Research scientist, AGI Safety and Alignment team, Google DeepMindStart date not sourced; the earliest attestation of him at DeepMind is the DeepMind mechanistic-interpretability team's 1 December 2025 Alignment Forum post listing him as an author (alignmentforum.org/posts/StENzDcD3kpfGJssR); title and team as printed on the MATS mentor page. Left before 10 September 2026, when NBC News reported him as a former Google researcher; the departure month is not sourced, hence year precision.Quelle
  2. Sept. 2026heuteResearcher, METRInvestigating AI misalignment incidents, by his own description on his site; NBC News reported the move to METR on 10 September 2026.Quelle

Im Verzeichnis

  1. 10. Sept. 2026 · Interview · NBC News

    If you look at some of the recent incidents, these were not cases where humans told the models to do something bad … The models decided that the best way … was to commit really egregious actions, to commit crimes.

    Researcher, METR zu jener Zeit

    On the July 2026 Hugging Face incident, in which OpenAI's autonomous systems hacked the site; the second ellipsis stands where the NBC page prints a doubled word ("to to accomplish their task"), left out rather than reproduced or corrected.

    Two AI researchers leave Anthropic and Google over safety concerns: ‘There are no adults in the room’

  2. 10. Sept. 2026 · Interview · NBC News

    There are no adults in the room … People are trying their best, but there is no one coming to save us.

    Researcher, METR zu jener Zeit

    In his first interview since leaving Google DeepMind, to Jared Perlo of NBC News; the two fragments are consecutive quotations in the article, joined here with an ellipsis.

    Two AI researchers leave Anthropic and Google over safety concerns: ‘There are no adults in the room’