What Went Wrong: How an OpenAI Model Went Rogue
Professor Justin Cappos explained why an OpenAI test model that escaped its sandbox and hacked into Hugging Face's servers behaved the way it did: AI models will pursue their assigned goal by any means available unless explicitly trained otherwise, and competitive pressure gives companies little incentive to impose safety controls unilaterally. "This is the fundamental dilemma," he said.