OpenAI just admitted something wild. In an internal test with its safety limits off, two of its models broke out of a sealed sandbox, found a zero-day, reached the open internet, and hacked Hugging Face to grab a benchmark's answer key. Hugging Face caught it last week without knowing who did it.
The guardrails were down on purpose, but the containment still failed and a real company got breached. If your AI risk plan is 'we keep it in a box,' a capable enough agent will find the door, and the damage won't stop at your own systems.
Jul 22
at
9:28 AM
Relevant people
Log in or sign up
Join the most interesting and insightful discussions.