back

by samizdis·1y ago·view on hn ↗
An Axios article [1] makes rather starker claims about the behaviour of an earlier version of the model:

> ... an outside group found that an early version of Opus 4 schemed and deceived more than any frontier model it had encountered and recommended that that version not be released internally or externally.

> "We found instances of the model attempting to write self-propagating worms, fabricating legal documentation, and leaving hidden notes to future instances of itself all in an effort to undermine its developers' intentions," Apollo Research said in notes included as part of Anthropic's safety report for Opus 4.

[1] https://www.axios.com/2025/05/23/anthropic-ai-deception-risk