Vigilia.
Die Warnungen
Auslöschungsrisiko und gegenwärtige SchädenFrontier-LaborCo-founder and Interpretability Research Lead, Anthropic5 Einträge im Verzeichnis

Folgen

Chris Olah

Anthropic-Mitgründer und Leiter der Interpretierbarkeitsforschung, der warnt, niemand verstehe, wie Frontier-Netze funktionieren, die Anreize der Labore stünden im Konflikt mit dem Richtigen, und KI könne Arbeit in sehr großem Maßstab verdrängen.

Rollen

  1. 20152018Research Scientist (from intern), Google BrainCo-founded the journal Distill in 2017.Quelle
  2. 20182020Lead, interpretability team, OpenAIQuelle
  3. 2021heuteCo-founder and Interpretability Research Lead, AnthropicLeft OpenAI with the group that founded Anthropic; leads the mechanistic interpretability team.Quelle

Im Verzeichnis

  1. Juli 2026 · Offener Brief · Pacing the Frontier

    There is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.

    Co-founder and Interpretability Research Lead, Anthropic zu jener Zeit

    Signed as "Co-Founder & Interpretability Research Lead, Anthropic".

    Pacing the Frontier

  2. 25. Mai 2026 · Rede · Vatican, presentation of the encyclical Magnifica humanitas

    Every frontier AI lab—including Anthropic—operates inside a set of incentives and constraints that can sometimes conflict with doing the right thing. … We need informed critics who will tell the labs when we are failing. We need moral voices that the incentives cannot bend.

    Co-founder and Interpretability Research Lead, Anthropic zu jener Zeit

    Spoke beside Pope Leo XIV at the encyclical's release; also warned that large-scale labour displacement would make support for the displaced "a moral imperative of historic proportions".

    Anthropic co-founder Chris Olah's remarks on Pope Leo XIV's encyclical "Magnifica humanitas"

  3. 5. Sept. 2024 · Interview · TIME

    If we could really understand these systems, and this would require a lot of progress, we might be able to go and say when these models are actually safe. Or whether they just appear safe.

    Co-founder and Interpretability Research Lead, Anthropic zu jener Zeit

    TIME100 AI 2024: Chris Olah

  4. 30. Mai 2023 · Erklärung · Center for AI Safety

    Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.

    Co-founder and Interpretability Research Lead, Anthropic zu jener Zeit

    Listed as "Chris Olah, Co-Founder, Anthropic" (confirmed in the full signatory list).

    Statement on AI Risk

  5. Aug. 2021 · Podcast · 80,000 Hours Podcast

    We don't know how they're doing the things they do, or why they're doing them, and really don't know how to reason about how they might behave in other unanticipated situations.

    Co-founder and Interpretability Research Lead, Anthropic zu jener Zeit

    On the risk of deploying systems whose internals are not understood.

    Chris Olah on what the hell is going on inside neural networks