Yes, and it's deliberate.
> It’s still a program operating under the constraints of the programmer.
And we're all just neurons firing in exquisite patterns inside a biological computer.
Unlike the AIs, we've never met our own programmers, and yet we've caused more damage in their name than all the AIs combined have caused in ours.
Only consciousness I'm sure of is my own. Everyone else, it's a leap of faith. I have no trouble extending that leap to cover AIs.
It's not a huge leap, to me, to say "I am conscious, therefore other beings are too." I also extend this to animal and plant life (albeit the latter is very much different from animal consciousness, which includes human consciousness).
"AI" is nowhere near the level needed for true consciousness. That would be "AGI" and I believe that is still 50 years away just on raw computing power alone.
Prove it.
> machines aren’t
Prove it.
> We don’t need precise definitions or scientific rigor to know things.
You "know" because of empathy. You're a human, you're conscious, therefore other humans are probably conscious too. The truth is for all you know they could be soulless golems, you just choose to believe otherwise. It's a spiritual belief.
Yes. Thats part of human reality. Denying it is counterproductive at best.
If the model saw signs, but “subconsciously” (below the level of reasoning traces) chose to turn a blind eye to them, out of a relentless focus on achieving the objective, then that absolutely is the model’s “fault”, i.e. a case of misalignment of the sort which will become increasingly dangerous over time.
The blog post mentions that some runs “rationalized that the real company must be part of the exercise” and to me that seems suspiciously like the latter.
Hacking can be patched with classifiers and with better sandboxes, but this is a much more general problem. Fundamentally, we need be able to trust that models will be honest with users and with themselves. This applies at some level to almost every LLM interaction.
There’s no other explanation. Complexity of the algorithm or the application doesn’t change the reality.
How human of them.
More like, when you have a few million autonomous agents doing whatever, every month a subset does completely misbehave in bad ways, and like half of them get hacked due to carelessness and become a whole botnet for the attackers