Vigilia.
Gli avvertimenti
Rischio di estinzioneLaboratorio di frontieraLead, Alignment Science team, Anthropic2 voci a registro

Seguire

Jan Leike

Ricercatore di allineamento che ha co-diretto il team di superallineamento di OpenAI, dimessosi nel maggio 2024 dicendo che la sicurezza era passata in secondo piano rispetto ai prodotti, oggi lavora in Anthropic, dove sostiene che supervisionare modelli sovrumani resta un problema irrisolto.

Ruoli

  1. 2021mag 2024Head of Alignment, OpenAICo-led the Superalignment project with Ilya Sutskever from June 2023 until his resignation.Fonte
  2. mag 2024oggiLead, Alignment Science team, AnthropicTeam focused on scalable oversight, weak-to-strong generalization and automated alignment research.Fonte

A registro

  1. 22 gen 2026 · Post · Musings on the Alignment Problem (Substack)

    Once our models are so capable that on many tasks we don't understand their actions anymore, a lot of current approaches won't work the same way.

    Lead, Alignment Science team, Anthropic all’epoca

    Alignment is not solved but it increasingly looks solvable

  2. 17 mag 2024 · Post · X

    safety culture and processes have taken a backseat to shiny products

    Lead, Alignment Science team, Anthropic all’epoca

    From the X thread announcing his resignation from OpenAI, as quoted by MIT Technology Review.

    Join me at EmTech Digital this week!