Anthropic admits AI 'not perfectly aligned' with human values(theguardian.com) |
Anthropic admits AI 'not perfectly aligned' with human values(theguardian.com) |
I mean basically. The hack the business events were result of training model for hacking, then testing its hacking abilities while not sandboxing it properly.