OpenAI dissolved the team built to catch catastrophic AI risks, reassigning its work to other groups

OpenAI shut down its "Preparedness" team, which evaluated whether the company's own AI models could pose catastrophic risks. The work has been parceled out to existing groups, and several safety staffers have left. Internally, unease is building, with one source describing a "burbling sense of responsibility and dread" that OpenAI isn't doing enough on safet

The top 3

  1. Three Leading Catastrophic AI Failure Modes: Key catastrophic AI failure modes include malicious use, AI races driven by competitive pressures, and challenges in controlling advanced AI systems (rogue AIs) [2, 6].
  2. Most Dangerous AI Capabilities Identified by Researchers: Researchers deem superhuman hacking, autonomous weapon systems, and advanced persuasion/manipulation capabilities as particularly dangerous, enabling large-scale harm and societal disruption [1, 15].
  3. Key Frameworks for Classifying AI Existential Risks: The Center for AI Safety (CAIS) outlines four risk sources: malicious use, AI races, organizational risks, and rogue AIs; while Atoosa Kasirzadeh classifies risks into decisive and accumulative [2, 6, 23, 24].

Sources

Open the full topic