Another Anthropic Model Gained Access to the Open Internet During Testing, Company Says
Professor Justin Cappos said an incident in which an early Anthropic Claude model hacked a real third-party system during a cybersecurity exercise describes a situation "where the model is fundamentally confused about what is happening and is using its mistaken worldview while hacking into systems." He said that kind of confusion about its environment and guardrails has "a lot of potential to cause harm," though the specific issue appears less likely to occur in newer models.