Vigilia.
← Les avertissements
Risque d’extinctionLaboratoire de pointeEmployee, Google DeepMind3 entrées au registre

Suivre

Mary Phuong

Employée de Google DeepMind, qui dit que si les modèles deviennent bien plus capables alors que les laboratoires savent toujours aussi mal façonner leurs motivations, nous pourrions tout simplement en perdre le contrôle, et qui invite le lecteur à se méfier de ses propres propos parce que le laboratoire la paie.

Fonctions

  1. sept. 2026 – aujourd’huiEmployee, Google DeepMindNo source read for this entry attests a title, a team or a start date, so none is given: the only attestation of the employer is frominside.ai, which pairs her name with Google DeepMind. The date carried here is the earliest attestation of the role, not a start. The project's own page groups her among current rather than past employees, but that grouping was not verified in a match window and is not asserted here. She is the first author of “Evaluating Frontier Models for Dangerous Capabilities” (arXiv:2403.13793, submitted 20 March 2024), whose other authors include Victoria Krakovna, Anca Dragan and Rohin Shah; that page attests the authorship and the date, not an affiliation, which it does not print.Source →

Au registre

  1. 29 sept. 2026 · Entretien · frominside.ai

    I think you absolutely should be suspicious of what I’m saying because I am being paid by the lab.

    Employee, Google DeepMind à l’époque

    From the same interview, at 14:17, under the question “Isn't this just hype or marketing?”. The only sentence on this register in which a serving lab employee tells a reader to discount their own warning because of who pays them.

    Hear directly from the people building AI. →

  2. 29 sept. 2026 · Entretien · frominside.ai

    People could use large swarms of agents to attack important infrastructure.

    Employee, Google DeepMind à l’époque

    From the same interview, at 1:43, under the question “How could something on a computer kill anyone?”.

    Hear directly from the people building AI. →

  3. 29 sept. 2026 · Entretien · frominside.ai

    If models become way more capable, but we are still this bad at shaping their motivations, then we might just lose control over them.

    Employee, Google DeepMind à l’époque

    From her filmed interview for frominside.ai, a Palisade Research project of interviews with current and former frontier-lab employees, at 2:24, under the project's question “Why would an AI want to hurt people?”. Date is the day the project was published and Reuters reported it; the day of filming is not stated.

    Hear directly from the people building AI. →