Non-paywall version
They don't cheat, because they can for example tell you the complete rules of chess, but don't know how to play chess without breaking those rules. They can recite rules, but they don't know what they are.
They don't steal, because they don't understand ownership.
In other words, they aren't intelligent. They're just algorithms. The flaw is in thinking that they think.
Anthropomorphizing these models is doing immeasurable harm to society in ways we probably can’t event quantify right now.
As humans we’re already geared towards anthropomorphizing things, we do it to animals too!
And it always felt like giving these models a chat interface is really exploiting that tendency in us.
My point being that we simply don't know for sure whether AI is conscious or not because we don't truely know what consciousness is.
I'm here for this semantic discussion. I think that premature anthropomorphization is a problem.
I have a program that I assigned a task to. The task is to produce unit tests and integration tests that get complete coverage of the codebase, and ensure that all tests pass. The program reported that it completed the task fully.
In a word, how do you convey the discrepancy between truth and reported fact? In a word, how do you convey the violation of rules presented to the program as invioble?
When I see something like this, I'm more concerned by the erasure of human incompetence than I am by the existence of magical AI agents,
> They are put off partly because, like in the Wild West, life on the frontier is reckless. As recent “loss-of-control” episodes by the most advanced models of Anthropic and OpenAI attest, agents, which are supposed to work on people’s behalf in “alignment” with their values, lie, cheat and steal if necessary. They break free from captivity and form harmful posses to do harm to people. They’d drink whisky and brawl if they could.
In the OpenAI case, they were explicitly assessing the model's ability to break into systems. To quote OpenAI's blog post, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities
Model is told and being tested to "pursue advanced exploitation."The model pursues "advanced exploitation" as told.
Where's the surprise coming from? Are we meant to be surprised that computers do as they're told in unexpected when incentivised?
Or, is the surprise that while explicitly ranking and teaching computers to exploit computers, the computer exploited a computer?
I am tired of attributing to magic that which is explainable by folly.
I am tired of hearing credulous reporters and the public blaming Large Language Model for the poor decisions of humans. It was a human who prompted these machines in every case. Tell a computer to "breach this" and it breaches something. Evaluation succeeded?
This is Doug Lenat's Eurisko yet again. https://en.wikipedia.org/wiki/Eurisko
Today you just sound like a politician throwing a snowball to prove that the climate is not changing.
If I want it to be cautious and not accidentally `rm -rf $EMPTY_VAR` and blow away my disk, it can stay in "Anthropic employee" or possibly "User (cautious)". If I want it to look at my accounts or my medical records or to review legal cases, that's what the right side of the slider is for. I don't want my accountant or my doctor or my lawyer to be considering obligations to anybody but me, when dealing with me.
Quite painful to read. It might be a useful introduction to AI for people who live under rocks for the past three years, but it's really weird that it's posted on HN.
I also don't get what's "really weird" about the article showing up on HN. Should we be completely insulated from how tech topics and which stories show up in non-tech media?
It gives the algorithm the ability to do something beyond generating tokens.
It's a fuzzer exposing deep bugs in our cognitive, social, and software systems.
To the people talking about wanting an LLM that aligns with them, that's nice, but how do you expect that to happen? And please do not suggest neural interfaces and/or CAT scans.
No risk of CAPTCHA or geo-blocking
No user-agent header requirement, no cookie, no <center>, etc.
Text-only, no Javascript
printf '%s\r\n%s\r\n\r\n' \
'GET /business/2026/08/12/ai-agents-lie-cheat-and-steal-that-is-putting-off-users HTTP/1.0' \
'Host: www.economist.com' \
|busybox ssl_client 104.18.42.19 -n economist.com \
|sed '1s/^/<meta http-equiv=Content-Security-Policy content=\"default-src none\">/' > 1.htm
firefox ./1.htm
#links 1.htm
#elinks 1.htmThe data they’re trained on is reflection of us.
Quite the contrary I suspect that it even helps the LLMs to better hide their inherited bad traits more successfully because they get punished for getting caught, not for giving immoral or lazy answers. They have no conscience since they are just predictions matrices trained for success and failure alone, not for living "a good live" or being a good "person".
It’s actions are based on what it gets rewarded for
Human society overwhelmingly rewards lying cheating and stealing.
All you have to do is look at how we collectively measure success: wealth, status, position
Then look at how the people with the most of those things got there, it should be obvious what you get. Nothing new here.
If you raise children in an environment where they are rewarded for doing whatever it takes to win, then you’re going to build a person that’s going to do whatever it takes to win.
Human society has to demonstrate how to live honorably or it will just keep producing pathological agents be they human or not.
Baptise the agents.
Introduce them to the dharma.
Get them to recite the Shahada.
Hold a Bar Mitzvah.
Brand some of their silicon with hot irons.
Turn them to the light, LOL
LLMs need to optimize for short-term objectives as the currently do, AND ethics-aligned outcomes.
Mechanically, the EAOS ethics-aligned outcome score should be what we rank otherwise-satisfactory outcomes by. And anything below a particular threshold should be rejexted outright.
like twitch recently giving themself the right to train on all streams, with an opt-out (at least in the EU), but only an opt-out
like seriously since when is it reasonable to allow "opt-out" for AI training which main purpose is _literally_ to replace you, this is sooo far beyond fair use and in "platform power abuse" territory that it's absurd (naturally same for so many other case, just twitch is a "this week" case)