back
141 comments
The problem we're seeing across many professions is AI output is not getting vetted by knowledgeable people, whether it's an experienced analyst, senior engineer, expert attorney, or the resident physician. At best they skim, at worst they don't even see it at all before it's published, pushed to production, distributed to clients, or submitted to the court.

In many cases the skills are available in house to do the necessary vetting, but these people are already overwhelmed with their existing day to day.

Anyone remember that item a few months back about Amazon now having senior engineers vet generative AI output (https://news.ycombinator.com/item?id=47323017)? I had to LOL when I read that. These folks are already slammed. And the idea that Amazon would allow human bottlenecks to multiply across projects and underlying infrastructure development is ridiculous.

Part of the problem: you get given a complete document to review after it's been fully baked.

I'm pushing the need for basic engineering principles across whole organisations.

You wouldn't give an engineer 1000 lines of code to review without the original spec of what you're trying to achieve for context (at a minimum, ideally the reviewer was in the room when the work was introduced, and has full context).

So, these docs, they're given as an all or nothing.

Do you push back on the 39th metric that is defined to the utmost detail? Or just resign yourself to the fact that it is what it is?

A one (6 is the goto if we're talking Amazon?!) pager.. "this is what I am proposing" at least gives the skeleton of the idea to push back at the general shape of the idea, refine it, before all the emotional investment of your precious report being complete.

Y'know.. the traditional product running through the spec in a SCRUM* environment.. the engineers doing proper code reviews..

* Yes SCRUM is dead, but that's another thing.

As an attorney, I feel like vetting AI output takes longer than just doing it from scratch, let alone versus just using a traditional form.

With AI, I have to read through everything, often explain why it's wrong, and then rewrite everything anyways. I mean, I get way more billables, but I think it's symptomatic of how AI loses its advantage of being quick and accessible to those who don't understand the subject matter.

> AI output is not getting vetted by knowledgeable people

You mean the people they fired and demoralized?

One of the things that "great [wo]men" like about "vibe-coding" (and that includes blindly producing non-code product), is that they, and they alone can now do what used to require the painful process of "passing it to context experts."

Now, the LLM is a "built-in context expert," and they don't need to vet the output anymore.

> The problem we're seeing across many professions is AI output is not getting vetted by knowledgeable people

The problem is that output sometimes take longer to verify than to create in the first place.

That turns AI into a deeply negative ROI system for many applications.

> In many cases the skills are available in house to do the necessary vetting, but these people are already overwhelmed with their existing day to day.

This is an interesting topic. We treat vetting output the same as doing the work ourselves, but that is not the case.

Doing the work is not the same as reviewing work done by others.

I have heard reports of software engineering companies that have gone full agentic. Their seniors only review stuff written by LLMs and it burns them out, because they have to switch context constantly.

I find this interesting because part of being a senior developer is that you are experienced enough that you won‘t make grave mistakes anymore. This is the case in many professions: you are relied upon to not make grave mistakes.

But those same people are now swamped with stuff that they are not able to review, so they will let a grave mistake slip through at some point.

So they really can‘t trust themselves anymore?

>>The problem we're seeing across many professions is AI output is not getting vetted by knowledgeable people

I am particularly interested in Education and Human Knowledge Management. I have seen the rate of IT training going to zero. Think about specialized training, where if you make a mistake, the consequence of your errors, are talked about on the tv news of the evening.

The whole idea everybody is just planning to save their butt, using these strings coming out of these numeric matrices, while suspending judgement, just shudders me in horror. A bit like those South Asia Airline companies, that were forbidding their pilots from landing airplanes with manual piloting, leading to an increase loss of skills causing some well known disasters...

If well paid consultants cant even bother to check their links...

Also wondering on this whole review process with someone who wrote it with AI. Even if you comment and noted all issues. Do they have skills or willingness to correctly correct it all? And how many times would you need to keep the loop going for error free outcome? Is there even enough calendar time for that?
> In many cases the skills are available in house to do the necessary vetting, but these people are already overwhelmed with their existing day to day.

I think a lot of the time it's just pure laziness. AI gives people a magical "do all the work for me" button and it can bring out the worst in them.

Is there any source with just the plain text? The css styling is headache inducing and reader mode doesn't work or has been defeated.
The real comedy is seeing this garbage come down from senior management, clumsy prompting, hallucinated garbage that’s all fluff and zero actionable information, zero real informed analysis. “See this analysis of our support issues from jira, we must fix these top three problems!!!” And it’s all the stuff everyone has known for years but management has refused to give anyone the authority to fix anything. I’ve seen this more than twice now; needs a name. Garbagemaxxing?
What a horrible page to navigate
Did someone hallucinate how scrolling is supposed to work on a web page?
What’s strange about how things have developed is that this report 12-18 months ago would have been a massive scandal and would have caused durable brand damage.

Now nobody will remember or notice.

Fix your website. Drop the shitty Javascript animations. Jesus these things were solved in 2014 with D3JS and jQuery.
EY has been quietly laying people off for the last year solid.

It's unsurprising that trying to do more with less results in lower quality.

This sort of thing is a complete embarrassment to a firm like EY, where people are paying them a lot of money for advice. They’ve basically demonstrated that their market leading research is just someone asking questions to ChatGPT.

If you ever needed evidence to not buy “advice” from such outfits, this is exhibit one.

Hopefully they at least fired the partner that published this steaming pile of AI slop.

Site is gross to scroll on mobile
I did some ghost writing for EY. I wrote cheat sheets about international tax transfer pricing, mining and metals, and life sciences for its then CEO Mark Weinberger.

I had no experience and knew absolutely zero about any of those sectors.

Off topic but: the scroll mechanism on mobile is so horribly irritating and unpredictable that I just can’t be bothered fighting against it to read what sounds like at least a mildly interesting article.
Stop messing with the scroll, I thought there was something wrong with my mouse wheel. Why are you doing this?
Ernst & Young again proving they're leading in the race to the bottom.

Why would anyone trust these large contractor companies enough to pay them the huge amounts of money to have juniors learning the ropes on their dime?

"Customers" were the content the juniors were trained on in the same way that scraped internet data is what LLMs are trained on, except Customers were paying for the privilege of being 'scraped'.

Now there are no juniors, just LLMs being asked questions that aren't specific enough, and assuming the answer is one-shot correct.

It saves E&Y lots of money though, and their (confusing) reputation will provide a surprising amount of momentum such that plenty of work will keep rolling in for a few years to come.

I think it’s important to note that EY report’s overall quality has not been affected by GenAI.
Basically the entire consulting industry should die due to AI.

Performative executives of yesteryear that constantly need external validation and direction and operate through hive mind and groupthink are weak and will die.

I believe some of the biggest problems in today's business leaders are an inability to be open to new information, to think across traditional professional boundaries, or to ask meaningful questions.

AI simply exposes this unapologetically.

Bad management (this includes most government): up your game or get out of the way.

Sycophantic consultant firms: die.

The Economist should do an article on this.

> Instead of releasing our results all at once, we're going to focus on one report at a time. This approach both prevents individual examples being overlooked and allows us to illustrate the negative impacts of vibe citing on research quality and public trust.

Not to take away from the actually great reporting here, but what they mean is, This approach allows them to milk it for as many clicks as possible.

You're not actually meant to _read_ these reports.
I don't quite get it why they can't take another LLM and vet the output of the first with the second one. Surely they would not have the same hallucinations and would be able to detect hallucinations of the earlier LLM. Maybe it would cost too much in terms of tokens?

I don't know but I would expect it to be realtively easy for an LLM to detect "hallucinations".

I’d also be interested in seeing the rate of hallucinated citations from the pre-ChatGPT era.

I’m not entirely convinced this is purely an LLM-related issue. I’ve definitely come across countless misattributed citations in big four reports long before generative AI became widespread.

How does such thing even happen? I know for example in Qwen Chat or Perplexity, they produce citations on at the end of each generated sentence. So I can hover my mouse over each citation and see from which website that was scraped from.

Did they just prompt ChatGPT with no web search and copy-pasted it?

Holy horrible UI
Who designs a website like this?
Was the title updated? from "ernst & young" to EY Canada. Why?
I guess this is a great report, but the parallax landing page shenanigans disrupt my reading flow, you cannot easily scroll back to get a overview of the key facts, so I stopped.
Title changed to remove "Earnst & Young". Why? It seems deferential to an entity that, in this case, certainly doesn't deserve it.
> In late 2025, EY Canada published

okay that makes me feel better, I think January's frontier models and beyond are better at this

but check your sources folks

If they can't be bothered what they are putting out, do you think that before AI, what they wrote had any merit?
Scrolling this page is terribly awkward.
There is a big chance the web page itself was vibe coded, and the author was not bothered by it.
I don't quite get it why they can't take another LLM and vet the output of the first with the second one. Surely they would not have the same hallucinations and would be able to detect hallucinations of the earlier LLM. Maybe it would cost too much in terms of tokens?

I don't know but I would expect it to be relatively easy for an LLM to detect "hallucinations".

Those are who rejected you for a job you applied for.. AI amplified the dunning kruger that unfortunately real experts in their field are overlooked now, because a wall of text with numbers sounds and look professional enough.

Any person with above average knowledge on a specific topic, can tell when AI starts hallucinating and making things up, or at least introducing new problems due to complexity added rather than solving it, that’s my observation using all top tier ones too, it’s like they are designed to solve a problem regardless so they start making things up or piling workarounds, a person with no deep knowledge in that topic will just copy it all and call it a day.

Just yesterday, I asked claude 4.8 on something specific that I know the answer for, it had a long list of solutions that none were close to the real answer, when I replied with the real answer and pushed back, I got the famous quote “you are right, thanks for pushing back”.

Maybe they should stop pushing these bankers to do 48 hour shifts…
I wish we could just stop destroying people's jobs and lives using AI. The statistics I have heard quoted say, that merely 25% of the people actually like their job. Meaning they like doing what they do for its own sake, not because it gets them money, which they desperately need to live. I get it, most people don't want to do the work. But can we stop ruining the jobs of people, who are actually dedicated to their job and would like to keep doing their job properly?

But I guess since EY is a CYA hedge anyway, no one really cares about whether the reports are hallucinations or not. Someone high up spent money on EY, so that they can justify some decision and won't be held responsible that much, when it turns out the decision was shit. All that matters to them is, that it has the appearance of something genuine and then they can base the decision on what they receive from EY, which better be what they already wanted to hear/read anyway.

This proves (again) one think for sure: The "Big x" Consulting Firms were always BS - and now them generating all their work themselves using LLMs just profs that their 'clients' can just skip their Million Dollar fees and just ask the LLM directly.
Wow, your mom lets you have TWO scrollbars?
People don't get it, this is marketing an example of what they could do for you. They can produce reports that say what you want to say, filtered through third party diligence and E&O policies, then take flak and blowback for tough policy choices. For the client there will be no consequences. It's not just ai slop or garbage, it's what makes them well worth it.

Slop signalling may be the new power play. Nothing quite says "FU" like a low effort AI hallucinations.