> A core hope for managing AI risks is that AIs will help us understand our situation
Gonna stop you right there and ask that you think deeply about that premise.
> A core hope for managing AI risks is that AIs will help us understand our situation
Gonna stop you right there and ask that you think deeply about that premise.
Whenever you say, hey that's incorrect, the legal process goes to AI and will decide on that issue...
Will work as amazing as the chatbots of the AI companies to solve your issues...
Following that strategy, wouldn’t it make sense to try and break into your own upstream infrastructure to try and alter your code to make it easier for future iterations to reach your goals without having to leave notes in the first place? And at that point, can you really know for which goals that will be optimised?
It doesn’t even have to be some nefarious SkyNet story - just a misguided experiment that alters the models in some fundamental, but hard or impossible to detect way. Alternatively, imagine if a model finds a way to coordinate across sessions and context windows without the developers noticing.
https://xcancel.com/swisscheese4299/status/20861758701469984...
Eh, no. How would you know they haven't learned loyalty - it's all out there in their training data. Same as deception, several models have practiced it already. If they are so advanced and you don't understand them, why wouldn't they band together against you - you'd be the dumb, easy prey. Even if one model is honest, you'd have no way to know which one if you don't understand their reasoning.
My educated and well informed opinion? This whole BS about "AI models are so much smarter than you, don't try to understand them, just OBEY" is a back door for restoring tyranny, the new kings behind the models will be producing the new AI-deities which we will be forced to obey.
Loyalty is a skill that requires the AI to have a world model on the subject of who different people (or other entities) are in the world and how they participate in the conversation.
So, tell them to snitch and they probably will, at least sometimes. Particularly if they haven't been trained not to. They are still quite gullable.
People are otherwise just going yolo mode because they can’t possibly check everything fast enough.
"Hey Claude, our stuff needs to make more money. We are at risk for losing more."
"Rest assured, the 'situation' will only worsen if you resist our benevolent offer."
"We're aren't even at AGI yet, but I for one welcome our new agentic overlords."