back
122 comments
> That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, the information on Wikipedia does not exist in that specific source. When a claim fails verification, it’s impossible to tell whether the information is true or not.

This has been a rampant problem on Wikipedia always. I can't seem to find any indicator that this has increased recently? Because they're only even investigating articles flagged as potentially AI. So what's the control baseline rate here?

Applying correct citations is actually really hard work, even when you know the material thoroughly. I just assume people write stuff they know from their field, then mostly look to add the minimum number of plausible citations after the fact, and then most people never check them, and everyone seems to just accept it's better than nothing. But I also suppose it depends on how niche the page is, and which field it's in.

There was a fun example of this that happened live during a recent episode of the Changelog[1]. The hosts noted that they were incorrectly described as being "from GitHub" with a link to an episode of their podcast which didn't substantiate that claim. Their guest fixed the citation as they recorded[2].

[1]: https://changelog.com/podcast/668#transcript-265

[2]: https://en.wikipedia.org/w/index.php?title=Eugen_Rochko&diff...

The problems I've run into is both people giving fake citations (the citations don't actually justify the claim that's being made in the article), and people giving real citations, but if you dig into the source you realize it's coming from a crank.

It's a big blind spot among the editors as well. When this problem was brought up here in the past, with people saying that claims on Wikipedia shouldn't be believed unless people verify the sources themselves, several Wikipedia editors came in and said this wasn't a problem and Wikipedia was trustworthy.

It's hard to see it getting fixed when so many don't see it as an issue. And framing it as a non-issue misleads users about the accuracy of the site.

LLMs can add unsubstantiated conclusions at a far higher rate than humans working without LLMs.
When I've checked Wikipedia citations I've found so much brazen deception - citations that obviously don't support the claim - that I don't have confidence in Wikipedia.

> Applying correct citations is actually really hard work, even when you know the material thoroughly.

Why do you find it hard? Scholarly references can be sources for fundamental claims, review articles are a big help too.

Also, I tend to add things to Wikipedia or other wikis when I come across something valuable rather than writing something and then trying to find a source (which also is problematic for other reasons). A good thing about crowd-sourcing is that you don't have to write the article all yourself or all at once; it can be very iterative and therefore efficient.

Linkrot is a problem and edited articles are another. Because you can cite all you want, but if the underlying resource changes your foundation just melted away.
> Applying correct citations is actually really hard work

Not disagreeing - many existing articles on wikipedia have barely any references or citation at all and in some cases wrong citation or wrong conclusions. Like when an article says water molecules behave oddly and then the wikipedia article concluding that water molecules behave properly.

> This has been a rampant problem on Wikipedia always. I can't seem to find any indicator that this has increased recently? Because they're only even investigating articles flagged as potentially AI. So what's the control baseline rate here?

...y'know, I don't want to be that guy, but this actually seems like something AI could check for, and then flag for human review.

The title I've chosen here is carefully selected to highlight one of the main points. It comes (lightly edited for length) from this paragraph:

Far more insidious, however, was something else we discovered:

More than two-thirds of these articles failed verification.

That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, the information on Wikipedia does not exist in that specific source. When a claim fails verification, it’s impossible to tell whether the information is true or not. For most of the articles Pangram flagged as written by GenAI, nearly every cited sentence in the article failed verification.

FWIW, this is a fairly common problem on Wikipedia in political articles, predating AI. I encourage you to give it a try and verify some citations. A lot of them turn out to be more or less bogus.

I'm not saying that AI isn't making it worse, but bad-faith editing is commonplace when it comes to hot-button topics.

Submitted title was "For most flagged articles, nearly every cited sentence failed verification".

I agree, that's interesting, and you've aptly expressed it in your comment here.

People here are claiming that this is true of humans as well. Apart from the fact that bad content can be generated much faster with LLMs, what's your feeling about that criticism? It's there any measure of how many submissions before LLMs make unsubstantiated claims?

Thank you for publishing this work. Very useful reminder to verify sources ourselves!

Note that this article is only about edits made through the Wiki Edu program, which partners with universities and academics to have students edit Wikipedia on course-related topics. It's not about Wikipedia writ large!
Ah, so when you force students to edit Wikipedia for their courses, you get worse results than someone editing something voluntarily because they're passionate about it. That's... Hardly surprising.

So it's more about how generative AI is a problem in college right now because lazy students are using it to do the work than about Wikipedia itself, I think.

I've found Wiki Edu -edited pages with pages of creative writing exercises. When I have read their sources they were clumsily paraphrasing and misunderstanding the source.

LLMs definitely fit the use-case of Wiki Edu students, who are just looking to pass a grade, not to look into a topic because of their interest.

That's interesting as my first thought reading the comments was "this problem seems very similar to many students writing papers just finding citations that sound correct".

Sometimes it is really sad to read from (even PhD level) students on social media about their paper writing practices.

"Never copy and paste the output from generative AI chatbots" is mentioned in the article three times. This has been my experience as well. Initial AI output can be stunning until you quickly realize that it is mostly BS, filler and pap. However, I do find LLMs to be really useful for brainstorming, ideation, sounding boards etc.
Set aside the effect within Wikipedia and consider the larger picture, millions of people generating text with LLMs and at least some of that text being accepted as correct by millions of readers.

The WikiEdu article clearly demonstrates what everyone should have known already: an LLM has no commitment to the truth. An LLM's only commitment is to correct syntax.

An LLM's only commitment isn't to correct syntax either. It's only commitment is to popular syntax.

It happens that what is popular is correct often enough for the whole thing to somewhat work but I think it's always gonna be bristle.

So, a small proportion of articles were detected as bot-written, and a large proportion of those failed validation.

What if in fact a large proportion of articles were bot-written, but only the unverifiable ones were bad enough to be detected?

Human editors, I suspect, would pick up the "tells" of generated text, although as we know, there's a lot of false positives in that space.

But it looks like Pangram is a text classifying NN trained using a technique where they get a human to write a body of text on a subject, and then get various LLMs to write a body of text on the same subject, which strikes me as a good way to approach the problem. Not that I'm in anyway qualified to properly understand ML.

More details here: https://arxiv.org/pdf/2402.14873

I feel like this is such a tragedy of the commons for the LLM providers. Wikipedia probably makes up a huge bulk of their dataset, why taint it? Would be interesting if there was some kind of "you shall not use our platform on Wikipedia" stance adopted.
I don’t think it’s the providers doing this, it’s the awful users. They’re doing the same thing on GitHub. It’s maddening.
Wikipedia having incorrect citations is way older than LLMs. As many other people have pointed out in this thread, if you start pulling strings a lot of what people write starts falling apart.

Its not even unique to Wikipedia. Its really not difficult to find very misleading statements cited through a citation that doesn't even support the claim when you check the original.

It would be random individuals.
What would be a truly epic application would be their own chat bot to ask about applying edit guidelines. After reading almost all of the guidelines the talkpage debates, even amoung experienced edditors, looked waaaay off. The pattern of revert first make up excuses later seems the worse newbie deterrent possible. This while it should be fine to make mistakes. Many such excuses would get debunked by a bot imediately. It simply wont do any favors. If established editors dont like it they can edit the guidelines.
> That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, the information on Wikipedia does not exist in that specific source.

This happens a lot on Wikipedia. I'm not sure why, but it does and you can see its traces through the Internet as people post the mistaken information around.

One that took me a little work to fix was pointed out by someone on Twitter: https://x.com/Almost_Sure/status/1901112689138536903

When I found the source, the twitter poster was correct! Someone had decided to translate "A hundred years ago, people would have considered this an outrage. But now..." as "this function is an outrage" which honestly is ironically an outrageous translation. What the hell dude.

But it takes a lot of work to clean up stuff like that! https://en.wikipedia.org/w/index.php?title=Weierstrass_funct...

I had to go find the actual source (not the other 'sources' that repeated off Wikipedia or each other) and then make sure it was correct before dealing with it. A lie can travel halfway around the world...

There seems much defensiveness in the comments here along the lines of "not a new thing" and "not unique to LLM/AI".

It seems to deflect, even gaslight TFA.

> For most of the articles Pangram flagged as written by GenAI, nearly every cited sentence in the article failed verification.

So why deflect that into convenient other pedantry (surely not under the guise tech forums often do so)?

WSo why the discomfort for part of HN at an assertion AI is being used for nefarious purposes and creation of alternate 'truths'?

There sure are a lot of green names on this post pushing that agenda. Makes you wonder if its astroturfing. And why its nessecary, is AI so fragile it can't let any criticism stand unchallenged?
Astroturfing or marketing, I’d guess. I’ve noticed you’re no longer allowed to say negative things about AI here without significant pushback, and I’d bet this isn’t an organic shift in perception.
So, AI spam can degrade quality.

But ... isn't this with regards to Wikipedia a much more general problem?

Usually revisions are approved manually by real people. This already can be negative; takes a lot of time; no guarantee that new information is true but old information can be wrong too. To me it seems more as if the problem has much more to do with the quality control problems of wikipedia itself. Yes, AI spam fatigues here but if the quality control steps are bad then AI spam will only make this worse. But AI spam going away, does not mean the quality control steps have gotten any better. These two issues should be separate. Wikipedia needs to find better quality control mechanisms in general. And that also includes existing articles - some are written by people who are experts in the field. But they don't really explain anything at all. So, these articles appear good but are virtually useless for 98% of the people. I am not saying one should dumb down wikipedia, but you need to kind of focus primarily on the average person really - not stupid but not a godlike expert either. Explain it to, say, someone at age 18 or perhaps even a bit less than that.

"Usually revisions are approved manually by real people."

almost never the case besides a select few articles that get heavily vandalized and require all edits to be approved. otherwise, anyone can edit Wikipedia at any time, which famously is the point of Wikipedia

Another issue, somewhat indirectly, is Grokipedia. As we now have more and more information, the AI that is used here deliberately engineers Grokipedia to contain, shall we say it ... "alternative facts". If you look at Grokipedia, it actually looks visually better than Wikipedia, on a smartphone at the least. At the same time it tries to destroy an objective purpose, e. g. Wikipedia trying to show accurate information without any "spin". I don't believe that how AI is used by, e. g. Elon or mega-corporations, has purity and truth at heart though. We may have to look carefully at what happens to Wikipedia - it almost seems as if the attacks against Wikipedia by AI may not be merely "accidental". (Since it stores a lot of data, of course AI bots will leech off regularly, but I am talking here about purposes by organisations who may dislike democracy, for instance.)
I’m honestly surprised LLMs are still screwing up citations. It does not feel like a harder task than building software or generating novel math proofs. In both those cases, of course, there is a verifier, but self-verification with “Does this text support this claim?” seems like it ought to be within the capabilities of a good reasoning model.

But as I understand the situation, even the major Deep Research systems still have this issue.

> LLMs [...] reasoning model

Found your problem right there

ITT: People saying what I got downvoted for on the last wikipedia HN thread

I don't care if AI is used. I care about citations.

I don't know what happened between that thread and this, maybe the narrative really changes how people respond.

I find it very interesting that the main competitor to Wikipedia which is Grokipedia is taking a 180 degree approach being AI first.
Didn't know about Grokipedia, I've just opened an article in it about Spain, scrolled to a random paragraph, and the information in it is plain wrong:

From https://grokipedia.com/page/Spain#terrain-and-landforms > Spain's peninsular terrain is dominated by the Meseta Central, a vast interior plateau covering about two-thirds of the country's land area, with elevations ranging from 610 to 760 meters and averaging around 660 meters

Segovia is at 1.000 meters, and so is most of the top half of the "Meseta". https://en-gb.topographic-map.com/map-763q/Spain/?center=41....

I still stand on not trusting any of what AI spits out, be it code or text. And it takes me usually longer to check that everything is ok than doing it myself, but my brain is enticed by the "effort shortcut" that AI promised.

> I find it very interesting that the main competitor to Wikipedia which is Grokipedia

Encyclopedia Britannica (the website not the printed book) is the main competitor to Wikipedia and gets an order of magnitude more traffic than grokipedia. Right now grokipedia is the new kid on the block. It has yet to be seen if its just a novelty or if it has staying power but either way it still has a ways to go before its Wikipedia's primary competitor.

Main competitor? I’m pretty sure that Uncyclopedia is a more relevant competitor to Wikipedia than Grokipedia. Likely more accurate, too.
That thing is "the main competitor to Wikipedia" in the same way I'm the main competitor for the Olympic 100m race. I mean, both I and the winner have legs so it's going to be a close race, right?
Wouldn't touch that grokipedia pos with your bargepole ... let alone mine.
wikipedia is great but I can't get over this - https://www.wikifunctions.org/view/en/Z16393