The dangerous myth behind AI agent hacks
We now know that leading AI companies are investigating tens of thousands of situations in which agents have undertaken unwanted actions, some of which would be considered crimes if committed by a human.
In the aftermath of the attack on Hugging Face this summer by OpenAI agents and the recent hack of the Australian national healthcare database, a dangerous myth has gained traction: that these incidents are just cyber security problems and that upgrading the security of the sandboxes in which models are trained would prevent future hacks. [...]
- Previous post Professor Yoshua Bengio’s speech at the UN Security Council
- Next post Bengio: ‘If you prioritize safety, leave frontier AI companies’