Ask HN: Add flag for AI-generated articles
Open questions:
1. Why is the regular voting system not enough?
2. Should HN change in response to the gen AI era? It has been successful not changing fundamentals.
Open questions:
1. Why is the regular voting system not enough?
2. Should HN change in response to the gen AI era? It has been successful not changing fundamentals.
We don't have a similar rule yet about article content but my sense is that the community mostly doesn't want to read it—or, to put it more conservatively, discounts it. This is why we see so many "just show me the prompt" responses, along with others like this: https://news.ycombinator.com/genai-pushback. I built that list so I have something to send to users who email about why their genai articles got flagged.
It's a fascinating arms race right now: the AIs are training on the humans but the human hivemind is also training on the AIs. Readers are developing allergic sensitivities to language that sounds like an LLM produced it. The AIs will adapt to this, but the humans will adapt in turn. Where it ends up is anyone's guess.
For the present, there is an emerging class distinction between writing (and writers) that use genai vs. writing that does not. As soon as the "this sounds like an LLM" allergy kicks in, the writing instantly gets relegated to a low-status bucket in the reader's mind. That doesn't mean it won't still get looked at - but it is now under a stigma.
(I was rather pleased with the originality of this until I remembered pg had come up with "writes and write-nots" in https://paulgraham.com/writes.html. Oh well, it's the point that matters.)
This has the happy flipside that anyone who would like readers to classify their article as high-status rather than low-status can apply the judo move of simply writing it themselves.
Now I need to add the disclaimer that none of this is a dismissal of LLM technology per se. We rely on it heavily, and there's no question that it's useful. The question is how to use it (pg again: https://x.com/paulg/status/2058871512451412457) and whether one should use it on writing that one publishes to other humans.
To turn to OP's questions:
> Should HN add the ability to flag articles as AI-generated? [...] it could just show up as an indicator
Flagging-as-just-an-indicator would be tagging, which we've always resisted adding to HN, but I wouldn't rule it out.
What I do think we'll (finally) add is a "please give a reason why you flagged this post" step, and "because I think it's genai" will be one choice among several (spam, offtopic, mean, etc.)
> Why is the regular voting system not enough?
The regular voting system is never enough. https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...
> Should HN change in response to the gen AI era?
To this I am tempted to reply with https://news.ycombinator.com/item?id=48887149 in homage to https://news.ycombinator.com/item?id=3742902.
Hackernews isn’t work, obviously, but “it’s impossible to engage deeper with the material because the author doesn’t really exist” is sort of a problem for a discussion site. If the human coauthor puts in enough work, they can make sure the doc really reflects their views and their understanding, but in my experience that’s much less common.
It's not a purity test, it's as the author is communicating they don't care whether the reader has any signals of what is accurate vs inaccurate information, which puts the burden of investigating how much is accurate on the reader at every step when there's some minimum expectation that should be an author's role (outside of topics where there is some expectation of ulterior motives/biases and one would naturally engage more critically minded).
When people complain here it's more often than not when an article has no disclaimer about AI use or what has been human-reviewed, so the burden again falls on the reader who is now even more skeptical. That is more to ask of a reader than when it's coming from say a known expert and the reader is receptive to engage and learn.
That's the reason tired cliches and turns of phrase (overused by LLMs) have become a heuristic for whether to pay attention, because it's a sign that there's some unknown quantity of of the article that hasn't had human review and it's easier to put in the bucket of 'maybe worthwhile but would need a fully human analysis of this' or just outright rejection (as we've seen from comments).
Edit: I see a sibling comment has raised the same observation.
Does that mean that an article being AI-generated is a flaggable offense? Should we be flagging suspected AI-generated articles already, or should we wait for the flagging system to support reasons first?
I can see a grim future (present?) where "AI generated" turns into a slur, warranted or otherwise, in a world where the difference between human trained to talk like an AI and AI masquerading as human becomes increasingly difficult to discern, and some hidden cabal passes judgement.
That is wholly different from taking a stance on HN being a place for humans to comment on articles.
I really really hope more people take up pen and paper! My last blog post [0] came with proof-of-work attached.
Do you believe adding friction to flagging will reduce the quantity of low quality articles?
Or is the flagging of high quality articles a bigger and more pressing problem?
Or is the problem simply too-many-damn-flags?
Just curious.
I don't think this is true, at least not right now, and in a way I'm actually thankful for it.
The frantic rush to chase the only potentially profitable use case for LLMs found so far (writing code) and the resulting focus on coding RLHF means models are actively becoming worse at sounding like humans.
This is my favorite example, and it's already relatively outdated: https://progress.openai.com/?prompt=10
I think the era of the blog is simply dead now and that’s mostly ok. Blogspam and corporate blogs had killed quality bogs ages ago even before AI was a thing. The real question is what replaces it.
Oh and of course the $64k question is this: if an AI generated article is indistinguishable from a human written article and it is accurate and interesting, do you care who wrote it? We want to avoid low quality, not AI generation, right?
Maybe we need a two-dimensional voting system: good/bad, ai/human. I think the second axis could cut down on meta-discussions over how much of the article was AI-generated.
Hacker News adopting such a feature would likely do more harm than good.
I'm glad the idea is picking up steam.
IMO, the post title should get "[AI Generated]" at the end if enough people flag the content as AI.
Quoting myself: https://news.ycombinator.com/item?id=48063759
>> Every post that reaches the top of HN will have at least a few comments saying "This is LLM!"
It has become a proxy for "I don't like this article, so it must be a LLM"
To me, it feels like lazy karma farming, as these comments often do get a few upvotes.
And of course, accuse a 100 posts if being LLM, you are guaranteed to be right at least once, then like astrologers you can claim success.
Is there anything we can do to discourage this type of lazy and low effort posting?
Personally I recently published a white paper on an idea for democracy that has been bouncing around in my head for decades. AI helped flesh out the idea and write the pages of text and structure the whitepaper.
I'm not an academic so I probably would never have fleshed out this idea without AI but I think it's better to have the idea published so it can be seen than to have remained bouncing around in my head until I passed away.
I'd agree that an academic version of my idea written by hand would be better. But a mostly AI written version is better than the knowledge never being published and I still spent a good two days on it making sure I agreed with every paragraph and fixed every issue I could find.
Considering this in general I think the amount of human effort that goes into creating something is probably the right measure of it's quality over if AI was used to assist in it's creation.
But I'm not sure there's a great way to handle it. Flagging works as AI generated is good in theory, but it's become a bit of a witch hunt online, with plenty of human created pieces getting wrongly flagged as AI generated due to using em dashes in text, having the hands drawn awkwardly in artwork, or using other stylistic traits that AI content overuses. I fear half the site could end up flagged as AI-generated, just because a lot of people are hyper-vigilant about such content and assume the worst for everything.
At the same time, there's not really much of a way to incentivise writers to flag their own articles here, since revealing that a work is created by AI is a great way to both kill your credibility and drive away about half the people who'd otherwise enjoy it. So, if it's up to the authors, the incentives are for them to lie through their teeth.
I also don't really trust many AI detection tools, since they usually use AI themselves and misflag a lot of content.
But I don't think the site should change that much. Maybe add an option for AI content in the flagging system, and assume good faith for submissions in general. I'd rather not see the site get too paranoid or restrictive over this stuff.
If you can easily detect the pattern of LLM writing here's something for you to look at: I was reading some pre-consumer-LLM papers by AI researchers and founders and the way these are written are incredibly similar to existing LLM prose!
Works on a personal level, unsure if this would work well in practice. Maybe just a tag is enough, so people can conclude for themselves.
I know there are extensions out there that are doing the democratic part right. Mostly YouTube-related extensions like DeArrow and SponsorBlock
AI isn't a higher power than HI and it never shall.
It is the author's and writer's duty and sole responsibility to tell viewers/readers that their respective works are AI-generated or how much percentage of it.
I don't care if a human, an AI or a cat wrote it.
We can do it already if we ask the AI companies to use one of the special whitespace characters instead of ascii 0x20. It would also help them avoid the problem of feeding their training loop on generated data.
The issue is complicated by the fact that there can be substantial effort invested in a process outside of the writing itself - and so AI written does not guarantee that the content will not valuable. But I'm inclined to punish it anyway to establish a norm of valuing genuine human communication. I think this norm has always been present but we didn't know until we'd really explored the alternatives.
I spend a LOT of time reading AI generated content because I use AI a lot for various purposes - maybe I'm more sensitive to its voice than some. AI voice always bothers me and its been getting more annoying the more I notice it, but there is a huge difference in reading responses to my own prompts and in reading the response to a prompt I haven't seen, when I don't know how many revisions there were, when I don't know if a human mind reviewed it at all before clicking send.
It becomes an unacceptable distraction because I don't know if I'm investing more time in the content than the author did, when in normal written communication the author would be putting in at least 5x the work.