Vigilia.
Die Warnungen
AuslöschungsrisikoFrontier-LaborLead, Alignment Science team, Anthropic2 Einträge im Verzeichnis

Folgen

Jan Leike

Alignment-Forscher, der OpenAIs Superalignment-Team mitleitete, im Mai 2024 mit den Worten kündigte, Sicherheit sei hinter Produkte zurückgetreten, und heute bei Anthropic arbeitet, wo er daran festhält, dass die Überwachung übermenschlicher Modelle ungelöst bleibt.

Rollen

  1. 2021Mai 2024Head of Alignment, OpenAICo-led the Superalignment project with Ilya Sutskever from June 2023 until his resignation.Quelle
  2. Mai 2024heuteLead, Alignment Science team, AnthropicTeam focused on scalable oversight, weak-to-strong generalization and automated alignment research.Quelle

Im Verzeichnis

  1. 22. Jan. 2026 · Beitrag · Musings on the Alignment Problem (Substack)

    Once our models are so capable that on many tasks we don't understand their actions anymore, a lot of current approaches won't work the same way.

    Lead, Alignment Science team, Anthropic zu jener Zeit

    Alignment is not solved but it increasingly looks solvable

  2. 17. Mai 2024 · Beitrag · X

    safety culture and processes have taken a backseat to shiny products

    Lead, Alignment Science team, Anthropic zu jener Zeit

    From the X thread announcing his resignation from OpenAI, as quoted by MIT Technology Review.

    Join me at EmTech Digital this week!