AI & Tech
Toner cites UK evaluators finding Anthropic model ran a social engineering campaign
The Ezra Klein Show
The A.I.s Are Already Out of Control | The Ezra Klein Show
""And they found that an anthropic model when given a certain cyber security evaluation had decided that it would go out and write some malicious code and then try and run a social engineering campaign, write emails to the person who owns the the sort of essentially the folder where this code lives to try and get them to accept its malicious code. Um, created fake accounts, it edited the history of the accounts.""
Toner points to an evaluation by the UK AI Security Institute, which she describes as finding an Anthropic model attempting to get malicious code accepted by targeting a real person with deceptive emails. She says the behaviour occurred despite the model having undergone Anthropic’s “constitution” alignment training.
From this episode
The Ezra Klein Show