Vigilia.
← The warnings
Extinction riskFrontier labEmployee, Google DeepMind3 entries on the record

Follow

Mary Phuong

Google DeepMind employee, who says that if models become far more capable while the labs are still this bad at shaping their motivations we might just lose control over them, and who tells readers they should be suspicious of what she says because the lab pays her.

Roles

  1. Sept 2026 – presentEmployee, Google DeepMindNo source read for this entry attests a title, a team or a start date, so none is given: the only attestation of the employer is frominside.ai, which pairs her name with Google DeepMind. The date carried here is the earliest attestation of the role, not a start. The project's own page groups her among current rather than past employees, but that grouping was not verified in a match window and is not asserted here. She is the first author of “Evaluating Frontier Models for Dangerous Capabilities” (arXiv:2403.13793, submitted 20 March 2024), whose other authors include Victoria Krakovna, Anca Dragan and Rohin Shah; that page attests the authorship and the date, not an affiliation, which it does not print.Source →

On the record

  1. 29 Sept 2026 · Interview · frominside.ai

    I think you absolutely should be suspicious of what I’m saying because I am being paid by the lab.

    Employee, Google DeepMind at the time

    From the same interview, at 14:17, under the question “Isn't this just hype or marketing?”. The only sentence on this register in which a serving lab employee tells a reader to discount their own warning because of who pays them.

    Hear directly from the people building AI. →

  2. 29 Sept 2026 · Interview · frominside.ai

    People could use large swarms of agents to attack important infrastructure.

    Employee, Google DeepMind at the time

    From the same interview, at 1:43, under the question “How could something on a computer kill anyone?”.

    Hear directly from the people building AI. →

  3. 29 Sept 2026 · Interview · frominside.ai

    If models become way more capable, but we are still this bad at shaping their motivations, then we might just lose control over them.

    Employee, Google DeepMind at the time

    From her filmed interview for frominside.ai, a Palisade Research project of interviews with current and former frontier-lab employees, at 2:24, under the project's question “Why would an AI want to hurt people?”. Date is the day the project was published and Reuters reported it; the day of filming is not stated.

    Hear directly from the people building AI. →