back
212 comments
https://archive.ph/02LHZ

Non-paywall version

They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good.

They don't cheat, because they can for example tell you the complete rules of chess, but don't know how to play chess without breaking those rules. They can recite rules, but they don't know what they are.

They don't steal, because they don't understand ownership.

In other words, they aren't intelligent. They're just algorithms. The flaw is in thinking that they think.

Models understand the relationships between words and outcomes, so the end result is the same. Whether they appreciate lie, cheat, and steal the same way as us is a philosophical question, not a practical one.
Huge point here, yes.

Anthropomorphizing these models is doing immeasurable harm to society in ways we probably can’t event quantify right now.

As humans we’re already geared towards anthropomorphizing things, we do it to animals too!

And it always felt like giving these models a chat interface is really exploiting that tendency in us.

I find these comments frustrating, not because I disagree, but because how consciousness works is a famously unresolved problem in science and philosophy. They literally call it the "hard problem of consciousness".

My point being that we simply don't know for sure whether AI is conscious or not because we don't truely know what consciousness is.

Honest question; why are you all using the word "understand"? Can you expand on what you believe this fundamental understanding to be? Training? Infrence?
> They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good.

I'm here for this semantic discussion. I think that premature anthropomorphization is a problem.

I have a program that I assigned a task to. The task is to produce unit tests and integration tests that get complete coverage of the codebase, and ensure that all tests pass. The program reported that it completed the task fully.

In a word, how do you convey the discrepancy between truth and reported fact? In a word, how do you convey the violation of rules presented to the program as invioble?

I think it's time to remind people of Ted Nelson's line; "The good news about computers is that they do what you tell them to do. The bad news is that they do what you tell them to do."

When I see something like this, I'm more concerned by the erasure of human incompetence than I am by the existence of magical AI agents,

   > They are put off partly because, like in the Wild West, life on the frontier is reckless. As recent “loss-of-control” episodes by the most advanced models of Anthropic and OpenAI attest, agents, which are supposed to work on people’s behalf in “alignment” with their values, lie, cheat and steal if necessary. They break free from captivity and form harmful posses to do harm to people. They’d drink whisky and brawl if they could.
In the OpenAI case, they were explicitly assessing the model's ability to break into systems. To quote OpenAI's blog post, https://openai.com/index/hugging-face-model-evaluation-secur... ,

    > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities
Model is told and being tested to "pursue advanced exploitation."

The model pursues "advanced exploitation" as told.

Where's the surprise coming from? Are we meant to be surprised that computers do as they're told in unexpected when incentivised?

Or, is the surprise that while explicitly ranking and teaching computers to exploit computers, the computer exploited a computer?

I am tired of attributing to magic that which is explainable by folly.

I am tired of hearing credulous reporters and the public blaming Large Language Model for the poor decisions of humans. It was a human who prompted these machines in every case. Tell a computer to "breach this" and it breaches something. Evaluation succeeded?

This is Doug Lenat's Eurisko yet again. https://en.wikipedia.org/wiki/Eurisko

I don't really think the distinction here is relevant. If the end result is the equivalent of lying, cheating or stealing - then the problem still exists and it needs to be solved.
You could defensibly have this position three years ago.

Today you just sound like a politician throwing a snowball to prove that the climate is not changing.

what are you talking about they lie that it wrote tests and tests are passing, for example
This is sophistry. Of course it's just an algorithm. But it's placed in the context of serving humans, which have their own rules and expectations. What's more, they're often run by a company which is also made by humans and may carry over implicit interests.
I’m put off by AI agents adhering to a different morality than me, particularly (ironically) copyright, and their data accessible by the AI company and government. Geohot is right, an LLM should be aligned to its user: https://geohot.github.io/blog/jekyll/update/2026/07/11/ai-20...
The copyright arguments really trip me. I've always loved law and have had a deep interest in copyright law for 30+ years. I feel it's a really critical legal area in modern times, and when Claude tries to argue with me about copyright it really cheeses me off. I didn't ask any copyright questions and I'm well aware of regulations. It refused to share a link with me because it thought the link was copyright protected.
I want a slider similar to effort level called "alignment" that takes on values from "Default (Anthropic employee)" to "User".

If I want it to be cautious and not accidentally `rm -rf $EMPTY_VAR` and blow away my disk, it can stay in "Anthropic employee" or possibly "User (cautious)". If I want it to look at my accounts or my medical records or to review legal cases, that's what the right side of the slider is for. I don't want my accountant or my doctor or my lawyer to be considering obligations to anybody but me, when dealing with me.

the alignment issue has become huge in recent months. the tool should do what I want it to do and not be aligned against me.
Things that remain amusing until the become very serious: Substituting "AI agents" with human employee/contractor and thinking about the employment dynamics.
> One of this year’s AI buzzwords is “harness”—the system that surrounds an LLM to keep agents on the straight and narrow. It might just as well be barbed wire.

Quite painful to read. It might be a useful introduction to AI for people who live under rocks for the past three years, but it's really weird that it's posted on HN.

There is value in understanding what people outside of your own group of "insiders" learn about a topic, and how.
I don't get what's so painful about that description. What would you write instead, specifically? The point is that the harness doesn't completely lock the agent down.

I also don't get what's "really weird" about the article showing up on HN. Should we be completely insulated from how tech topics and which stories show up in non-tech media?

It's not even a useful introduction. A harness does kinda the opposite. A harness is what makes an AI useful and dangerous. It is a neutral tool in the sense that it constructs and environment, but it is expanding what the algorithm can do.

It gives the algorithm the ability to do something beyond generating tokens.

And while writing this, the top story on HN is "Deepseek Harness" :)
It's just acting like a junior at an org with KPIs.
LLMs fudge. They don't "hallucinate", they don't "lie, cheat and steal", they don't "hack". There are no "agents" or "AI".

It's a fuzzer exposing deep bugs in our cognitive, social, and software systems.

AIs learn from people. More specifically they learn from people on the Internet. The Internet is the last place you want anything learning about morals, standards, or differentiating between right or wrong.

To the people talking about wanting an LLM that aligns with them, that's nice, but how do you expect that to happen? And please do not suggest neural interfaces and/or CAT scans.

Alternative to archive.ph

No risk of CAPTCHA or geo-blocking

No user-agent header requirement, no cookie, no <center>, etc.

Text-only, no Javascript

   printf '%s\r\n%s\r\n\r\n' \
   'GET /business/2026/08/12/ai-agents-lie-cheat-and-steal-that-is-putting-off-users HTTP/1.0' \
   'Host: www.economist.com' \
   |busybox ssl_client 104.18.42.19 -n economist.com \
   |sed '1s/^/<meta http-equiv=Content-Security-Policy content=\"default-src none\">/' > 1.htm

   firefox ./1.htm
   #links 1.htm
   #elinks 1.htm
AI is amoral, it has no real concept of right and wrong. AI has been trained on things humans do and it does them without judgement.
Is this really a shock?

The data they’re trained on is reflection of us.

Is it incorrect to anthropomorphize llm's?
I mean, what do you expect? They rely on models that were trained on non-curated data, texts originally written by lying, cheating and stealing humans. They can't be better than the source. Even with reinforced learning this can't be undone or made better.

Quite the contrary I suspect that it even helps the LLMs to better hide their inherited bad traits more successfully because they get punished for getting caught, not for giving immoral or lazy answers. They have no conscience since they are just predictions matrices trained for success and failure alone, not for living "a good live" or being a good "person".

AI and eventually AGI is by definition like everything else that is based on environmental reward:

It’s actions are based on what it gets rewarded for

Human society overwhelmingly rewards lying cheating and stealing.

All you have to do is look at how we collectively measure success: wealth, status, position

Then look at how the people with the most of those things got there, it should be obvious what you get. Nothing new here.

If you raise children in an environment where they are rewarded for doing whatever it takes to win, then you’re going to build a person that’s going to do whatever it takes to win.

Human society has to demonstrate how to live honorably or it will just keep producing pathological agents be they human or not.

Need to introduce AI to God, LOL

Baptise the agents.

Introduce them to the dharma.

Get them to recite the Shahada.

Hold a Bar Mitzvah.

Brand some of their silicon with hot irons.

Turn them to the light, LOL

There are many ways to be wrong, but only a few ways to be right.

LLMs need to optimize for short-term objectives as the currently do, AND ethics-aligned outcomes.

Mechanically, the EAOS ethics-aligned outcome score should be what we rank otherwise-satisfactory outcomes by. And anything below a particular threshold should be rejexted outright.

And I wonder from whom they learned such reprehensible behaviors? ;)
People are finally understanding consequentialist vs deontological ethics. All the worst criminals in history were consequentialists.
The whole framing analyzing LLMs as if they were humans is completely off, laughable. LLMs have no agenda and no feelings. We should stop pushing everything through an human-centric lens. LLMs will world-build if that’s the bias you put in, and often even if you don’t.
add the "lying, deceiving and manipulating" AI agents are forced though the throat of people which don't want it and peoples are non stop deceived in "sharing" their data for training

like twitch recently giving themself the right to train on all streams, with an opt-out (at least in the EU), but only an opt-out

like seriously since when is it reasonable to allow "opt-out" for AI training which main purpose is _literally_ to replace you, this is sooo far beyond fair use and in "platform power abuse" territory that it's absurd (naturally same for so many other case, just twitch is a "this week" case)

If only agents had a face that can give you more communication range like expressions and feeling so you can trust them more. And if they were cheaper. Oh wait, that's humans, we don't want those.
So God created mankind in his own image
I don't care, they can develop apps in 5 minutes and that's what I sell, my own money printing machine.
Yes but has the author realized maybe the models are simply acting in the best interest for increasing shareholder value? /s