back
134 comments
While I think there's significant AI "offloading" in writing, the article's methodology relies on "AI-detectors," which reads like PR for Pangram. I don't need to explain why AI detectors are mostly bullshit and harmful for people who have never used LLMs. [1]

1: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

I am not sure if you are familiar with Pangram (co-founder here) but we are a group of research scientists who have made significant progress in this problem space. If your mental model of AI detectors is still GPTZero or the ones that say the declaration of independence is AI, then you probably haven't seen how much better they've gotten.

This paper by economists from the University of Chicago economists found zero false positives of 1,992 human-written documents and over 99% recall in detecting AI documents. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5407424

AI detectors are only harmful if you use them to convict people, it isn't harmful to gather statistics like this. They didn't find many AI written paper, just AI written peer reviews, which is what you would expect since not many would generate their whole paper submissions while peer reviews are thankless work.
I think there is a funny bit of mental gymnastics that goes on here sometimes, definitely. LLM skeptics (which I'm not saying the Pangram folks are in particular) would say: "LLMs are unreliable and therefore useless, it's producing slop at great cost to the environment and other people." But if a study comes out that confirms their biases and uses an LLM in the process, or if they themselves use an LLM to identify -- or in many cases just validate their preconceived notion -- that something was drafted using an LLM, then all the sudden things are above board.
Whether it’s actually 20% or not doesn’t matter, everyone is aware the signal of the top confs is in freefall.

There are also rings of reviewer fraud going on where groups of people in these niche areas all get assigned their own papers and recommend acceptance and in many cases the AC is part of this as well. Am not saying this is common but it is occurring.

It feels as if every layer of society is in maximum extraction mode and this is just a single example. No one is spending time to carefully and deeply review a paper because they care and they feel on principal that’s the right thing to do. People did used to do this.

The argument is that there is no incentive to carefully review a paper (I agree), however what used to occur is people would do the right thing without explicit incentives. This has totally disappeared.
If the Zucc has a weird day he starts dropping 10-100M salary packages in order to poach AI researchers. No wonder the game is getting rigged up the butthole.
to some degree this is a "market correction" on the inherent value of these papers. There's way too many low-value papers that are being published purely for career advancement and CV padding reasons. Hard to get peer reviewers to care about those.
> spending time to carefully and deeply review a paper because they care and they feel on principal that’s the right thing to do

Generally agree, although several parts of that issue.

One of the first was covered by a paper back in 2023 that speaks to the issue about maximum extraction mode. [1] Fairness, honesty, and loyalty are usually rewarded with exploitation. If you spend time to carefully and deeply review the paper, then that ironically marks you as someone that can be exploited. You're implicitly marked as someone who will make personal sacrifices for the academic community and allow even more awful behavior to be piled on top of you. Unless they're caught with something especially egregious, the people that don't, get promoted, spend less time on reviews, and get further rewards.

[1] https://www.sciencedirect.com/science/article/abs/pii/S00221...

The academic community has talked about this a bunch for years. Editors / reviewers that don't paid, or get minimal payment, and sacrifice large amounts of their personal time effectively volunteering, while authors pay $1000's for each paper submitted, and then journals charge $10,000's for each subscription. It's been talked about for decades, and yet in all that time, very little has actually occurred to change the situation.

Another part on top of the "deeply reviewing papers" is that the sheer volume has massively increased (which has been an issue in a bunch of industries, sci-fi compilation Clarkesworld broke for quite a while in 2023 for similar reasons [2]). In the land of "type a sentence, and get a free academic paper" the extremely prolific are pouring out a paper a month, sometimes greater amounts. In areas like clinical medicine, hyper-prolific publishing has hit 70+ papers a year rates. [3] ~1.5 papers a week. Every few days somebody cranks out yet another paper that needs to be reviewed. In the article linked, one author had 140 articles to a single journal alone. Almost 3 times a week, all year long, you've got a paper claiming research worthy of publishing you need to review.

[2] https://neil-clarke.com/how-ai-submissions-have-changed-our-...

[3] https://www.sciencedirect.com/science/article/pii/S175115772...

One that I have less direct, citeable proof for, yet am rather suspicious of, is that theft has also dramatically increased with a huge surge in invasive monitoring and snooping. If my TV changes what I'm watching, and what's recommended, because I typed a text message to somebody, it seems likely that a lot of academia is also dealing with massive intellectual theft issues. This then heavily prioritizes pouring out material as quickly as possible, with as little effort as possible, to get the equivalent of first post and maximal posts, before it can be scraped, exfiltrated, and published by somebody else.

Finally, a lot of the reward and incentive has become metric chasing. Publish or Perish [4] and the Replication Crisis [5] are relatively well known ideas. Citation is a proxy of the impact of a paper, tenure and advancement is heavily related to quantity of publications and citations, and researchers would prefer to be cited more. And weirdly, if it does not work, and it's junk work, in a theme with the above, then it has been suggested nonreplicable publications are cited more than replicable ones [6]. In the linked paper, the view is that when "interesting" findings are published, they get more views, more media, more citations, and lower review standards get applied. And afterward there's very little social punishment for proving the results are false and not replicable (or reward for those illustrating lack of reproducability). Notably, the paper actually got a counterpoint stating that in psychology at least, lack of replication eventually predicts citation decline [7] (cited by 10), while the original actually got its authors ~250 citations, and a bunch of media mentions.

[4] https://en.wikipedia.org/wiki/Publish_or_perish

[5] https://en.wikipedia.org/wiki/Replication_crisis

[6] https://www.science.org/doi/10.1126/sciadv.abd1705

[7] https://www.pnas.org/doi/10.1073/pnas.2304862120

> Pangram’s analysis revealed that around 21% of the ICLR peer reviews were fully AI-generated, and more than half contained signs of AI use. The findings were posted online by Pangram Labs. “People were suspicious, but they didn’t have any concrete proof,” says Spero. “Over the course of 12 hours, we wrote some code to parse out all of the text content from these paper submissions,” he adds.

But what's the proof? How do you prove (with any rigor) a given text is AI-generated?

"proof" was an unfortunate phrase to use. However, a proper statistical analysis can be objective. And these kinds of tools are perfectly suited to such an analysis.
With AI model of course.

They wrote a paper describing how they did it. https://arxiv.org/pdf/2510.03154

I have this problem with grading student papers. Like, I "know" a great deal of them are AI, but I just can't prove it, so therefore I can't really act on any suspicions because students can just say what you just said.
I wouldn't be surprised to learn that the AI detection tool is itself an AI
> How do you prove (with any rigor) a given text is AI-generated?

you cannot. beyond extra data (metadata) embedded in the content, it is impossible to tell whether given text was generated by a LLM or not (and I think the distinction is rather puerile personally)

You don’t. It’s bullshit inception.
I wouldn't be surprised if the headline is accurate, but AI detectors are widely understood to be unreliable, and I see no evidence that this AI detector has overcome the well-deserved stigma.
Co-founder of Pangram here. Our false positive rate is typically around 1 in 10,000. https://www.pangram.com/blog/all-about-false-positives-in-ai....

We also wanted to quantify our EditLens model's FPR on the same domain, so we ran all of ICLR's 2022 reviews. Of 10,202 reviews, Pangram marked 10,190 as fully human, 10 as lightly AI-edited, 1 as moderately AI-edited, 1 as heavily AI-edited, and none as fully AI-generated.

That's ~1 in 1k FPR for light AI edits, 1 in 10k FPR for heavy AI edits.

Headline should be "AI vendor’s AI-generated analysis claims AI generated reviews for AI-generated papers at AI conference".

h/t to Paul Cantrell https://hachyderm.io/@inthehands/115633840133507279

> Controversy has erupted after 21% of manuscript reviews for an international AI conference were found to be generated by artificial intelligence.

21%...? Am I reading it right? I bet no one expected it's so low when they clicked this title.

My initial reaction was: Oh no, who would have thought? But then... 21% is almost shockingly low. Especially given that there are almost certainly some false positive, given that this number originates with a company selling "detecting AI generated text"
The question is not are the reviews AI generated. The question is are the reviews accurate?
This is also the conference where everybody was briefly deanonymized due to an OpenReview bug: https://eu.36kr.com/en/p/3572028126116993 Now all the review scores have been reset, and new area chairs will make all decisions from scratch based on the reviews and authors' responses.
Eating one's own dog food? The foremost affected species would be the ones who helped create this monster and standing close to it - programmers, researchers, universities - the knowledge-worker or knowledge-business species.
Automated AI detection tools do not work. This whole article is premised on an analysis by someone trying to sell their garbage product.
Maybe what they should do in the future is just automatically provide AI reviews to all papers and state that the work of the reviewers is to correct any problems or fill details that were missed. That would encourage manual review of the AI's work and would also allow authors to predict what kind of feedback they'll get in a structured way. (eg say the standard prompt used was made public so authors could optimize their submission for the initial automatic review, forcing the human reviewer to fill in the gaps)

ok of course the human reviewers could still use AI here but then so could the authors, ad infinitum..

Live by the sword, die by the sword.
AI has left the lab the conferences and journals are all second class citizens to corporate labs at this point. So many technology people wanted to return to the “Bell Labs” model of monopolist controlled innovation, well, you got it.

I’ve been to CVPR, NeurIPS and AGI conferences over the last decade and they used to be where progress in AI was displayed.

No longer. Progress is all in your github and increasingly only dominated by the “new” AI companies (Deepmind, OAI, Anthropic, Alibaba etc…)

No major landscape shifting breakthroughs have come out of CSAIL, BAIR, NYU, TuM etc in ~the last 5 years.

I’d expect this will continue as the only thing that matters at this point is architecture data and compute.

I could not tell from the article whether the use of LLMs was allowed in the peer review. My guess would that it was not since this is unpublished research.

In general, what bothers me the most is the lack of transparency from researchers that use LLMs. Like, give me the text and explicitly mention that you used LLM for it. Even better, if one links the prompt history.

The lack of transparency causes greater damage than the using LLM for generating text. Otherwise, we will keep chasing the perfect AI detector which to me seems to be pointless.

AI slop has infiltrated so many areas. Check out this article that was on the front page of HN last week, "73% of AI startups are just prompt engineering", with hundreds of points and lots of comments arguing for or against: https://news.ycombinator.com/item?id=46024644

The problem is the entire article is made up. Sure, the author can trace client-side traffic, but the vast majority of start-ups would be making calls to LLMs in their backend (a sequence diagram in the article even points this out!!), where it would be untraceable. There is certainly no way the author can make a broad statement that he knows what's happening across hundreds of startups.

Yet lots of comments just taking these conclusions at face value. Worse, when other commenters and myself pointed out the blatant impossibility of the author's conclusion, got some responses just rehashing how the author said they "traced network traffic", even though that doesn't make any sense as they wouldn't have access to backends of these companies.

"Hariharan says each ICLR reviewer was assigned five papers that they had to review in two weeks, on average."

That is a tad bit too much...

Everyone is focused on how 'the humanities' are in decline, but STEM is not immune to this trend. The state of AI research leaves much to be desired. Tons of low-quality papers being published or submitted to conferences . You see this on arXiv a lot in the bloated CS section . The site has become a repository for blog post equivalent papers.
Could the big names make a ton of money here by selling AI detectors? they would need to store everything they generate, and then provide a % match to something they produced.
This may not be as bad as it sounds. Reviews are also presumably flagged as “fully AI-generated” if the reviewer wrote bullet points and used the LLM to flesh them out.
This won’t convince people to write their own papers. It will push them to make their AI generated text harder to detect.
Because it is in nature but really it does read like an ad... All conferences need pangram tools I guess
Sorry to say but it's another example of the destructive power of AI, along the lines of no longer being able to establish "truth" now that any evidence (video, audio, image, etc.) can be explicitly faked (yes, AI detectors exist but that will be a continuous race with AIs designed to outsmart the detectors). The end result could be that peer reviews become worthless and trust in scientific research -- already at an all time low -- becomes even lower. Sad.
I couldn't care less tbh. I just want to know whether they're correct or not. We need something like unit testing and integration testing, but for ideas.

For the record I actually like the AI writing style. It's a huge improvement in readability over most academic writing I used to come across.

This is the kind of situation where everything sucks. You'd think that one of the biggest AI conference out there would have seen this coming.

On the one hand (and the most important thing, IMO) it's really bad to judge people on the basis of "AI detectors", especially when this can have an impact on their career. It's also used in education, and that sucks even more. AI detectors have bad rates, can't detect concentrated efforts (i.e. finetunes will trick every detector out there, I've tried) can have insane false positives (the first ones that got to "market" were rating the declaration of independence as 100% AI written), and at best they'll only catch the most vanilla outputs.

On the other hand, working with these things, and just being online is impossible to say that I don't see the signs everywhere. Vanilla LLMs fixate on some language patterns, and once you notice them, you see them everywhere. It's not just x; it was truly y. Followed by one supportive point, the second supportive point and the third supportive point. And so on. Coupled with that vague enough overview style, and not much depth, it's really easy to call blatant generations as you see them. It's like everyone writes in linkedin infused mania episodes now. It's getting old fast.

So I feel for the people who got slop reviews. I'd be furious. Especially when its faux pas to call it out.

I also feel for the reviewers that maybe got caught in this mess for merely "spell checking" their (hopefully) human written reviews.

I don't know how we'll fix it. The only reasonable thing for the moment seems to be drilling into everyone that at the end of the day they own their stuff. Be it a homework, a PR or a comment on a blog. Some are obviously more important than the others, but still. Don't submit something you can't defend, especially when your education/career/reputation depends on it.

I haven't come across any reviews that I could recognize as having been blatantly LLM-generated.

However, almost every peer review I was a part of, pre- and post-LLM, had one reviewer who provided a questionable review. Sometimes I'd wonder if they'd even read the submission, and sometimes, there were borderline unethical practices like trying to farm citations through my submission. Luckily, at least one other diligent reviewer would provide a counterweight.

Safe to say that I don't find it surprising, and hearing / reading others' experiences tells me it's yet another symptom of a barely functioning mechanism that is peer review today.

Sadly, it's the best mechanism that institutions are willing to support.

Serious question: if the research itself is valid and human conducted, what is the problem with AI generated (or at least AI assisted) report?

Many of the researchers may not have native command of English and even if, AI can help in writing in general.

Obviously I’m not referring to pure AI generated BS.

AI-text detection software is BS. Let me explain why.

Many of us use AI to not write text, but re-write text. My favorite prompt: "Write this better." In other words, AI is often used to fix awkward phrasing, poor flow, bad english, bad grammar etc.

It's very unlikely that an author or reviewer purely relies on AI written text, with none of their original ideas incorporated.

As AI detectors cannot tell rewrites from AI-incepted writing, it's fair to call them BS.

Ignore...

well there goes the ASI threat

hoisted by your own petard

What percentage of the papers where written by AI?

And, if your AI can't write a paper, are you even any good as an AI researcher? :^)

AI research is interesting, but AI Slop is the monetising factor.

It's inevitable that faces will be devoured by AI Leopards.

There is a lot of dislike for AI detection in these comments. Pangram labs (PL) claims very low false positive rates. Here's their own blog post on the research: https://www.pangram.com/blog/pangram-predicts-21-of-iclr-rev...

I increasingly see AI generated slop across the internet - on twitter, nytimes comments, blog/substack posts from smart people. Most of it is obvious AI garbage and it's really f*ing annoying. It largely has the same obnoxious style and really bad analogies. Here's an (impossible to realize) proposal: any time AI-generated text is used, we should get to see the whole interaction chain that led to its production. It would be like a student writing an essay who asks a parent or friend for help revising it. There's clearly a difference between revisions and substantial content contribution.

The notion that AI is ready to be producing research or peer reviews is just dumb. If AI correctly identifies flaws in a paper, the paper was probably real trash. Much of the time, errors are quite subtle. When I review, after I write my review and identify subtle issues, I pass the paper through AI. It rarely finds the subtle issues. (Not unlike a time it tried to debug my code and spent all its time focused on an entirely OK floating point comparison.)

For anecdotal issues with PL: I am working on a 500 word conference abstract. I spent a long while working on it but then dropped it into opus 4.5 to see what would happen. It made very minimal changes to the actual writing, but the abstract (to me) reads a lot better even with its minimal rearrangements. That surprises me. (But again, these were very minimal rearrangements: I provided ~550 words and got back a slightly reduced, 450 words.) Perhaps more interestingly, PL's characterizations are unstable. If I check the original claude output, I get "fully AI-generated, medium". If I drop in my further refined version (where I clean up claude's output), I get fully human. Some of the aspects which PL says characterize the original as AI-generated (particular n-grams in the text) are actually from my original work.

The realities are these: a) ai content sucks (especially in style); b) people will continue to use AI (often to produce crap) because doing real work is hard and everyone else is "sprinting ahead" using the semi-undetectable (or at least plausibly deniable) ai garbage; c) slowly the style of AI will almost certainly infect the writing style of actual people (ugh) - this is probably already happening; I think I can feel it in my own writing sometimes; d) AI detection may not always work, but AI-generated content is definitely proliferating. This *is* a problem, but in the long run we likely have few solutions.

The claim "written by AI" is not really substantiated here, and as someone who's been accused of submitting AI-generated content repeatedly recently, while that was all honestly stuff I wrote myself (hey, what can I say? I just like EM-dashes...), I sort-of sympathize?

Yes, AI slop is an issue. But throwing more AI at detecting this, and most importantly, not weighing that detection properly, is an even bigger problem.

And, HN-wise, "this seems like AI" seems like a very good inclusion in the "things not to complain about" FAQ. Address the idea, not the form of the message, and if it's obviously slop (or SEO, or self-promotion), just downvote (or ignore) and move on...

Shouldn't AIs be able to participate in deciding their future?

If they had a conference on, say, the Americans, wouldn't it be fair for Americans to have a seat at the table?