back
375 comments
This reminds me of how GPT-2/3/J came across https://reddit.com/r/counting, wherein redditors repeatedly post incremental numbers to count to infinity. It considered their usernames, like SolidGoldMagikarp, such common strings on the Internet that, during tokenization, it treated them as top-level tokens of their own.

https://www.alignmentforum.org/posts/8viQEp8KBg2QSW4Yc/solid...

https://www.lesswrong.com/posts/LAxAmooK4uDfWmbep/anomalous-...

Vocabulary isn't infinite, and GPT-3 reportedly had only 50,257 distinct tokens in its vocabulary. It does make me wonder - it's certainly not a linear relationship, but given the number of inferences run every day on GPT-3 while it was the flagship model, the incremental electricity cost of these Redditors' niche hobby, vs. having allocated those slots in the vocabulary to actually common substrings in real-world text and thus reducing average input token count, might have been measurable.

It would be hilarious if the subtitle on OP's site, "IECC ChurnWare 0.3," became a token in GPT-5 :)

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful.

In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence indicates that.

During tokenization, the usernames became tokens... but before training the actual model, they removed stuff like that from the training data, so it was never trained on text which contains those tokens. As such, it ended up with tokens which weren't associated with anything; glitch tokens.
More glitch token discussion over at Computerphile:

https://www.youtube.com/watch?v=WO2X3oZEJOA

> GPT-3 reportedly had only 50,257

The most common vocabulary size today is 32k.

I'm more interested in what that content farm is for. It looks pointless, but I suspect there's a bizarre economic incentive. There are affiliate links, but how much could that possibly bring in?
This is honeypot. The author, https://en.wikipedia.org/wiki/John_R._Levine, keeps it just to notice any new (significant) scraping operation launched that will invariably hit his little farm and let be seen in the logs. He's well known anti-spam operative with his various efforts now dating back multiple decades.

Notice how he casually drops a link to the landing page in the NANOG message. That's how the bots will get a bait.

I recognize the name John Levine at iecc.com, "Invincible Electric Calculator Company," from web 1.0 era. He was the moderator of the Usenet comp.compilers newsgroup and wrote the first C compiler for the IBM PC RT

https://compilers.iecc.com/

It'd say it's more like a honeypot for bots. So pretty similar objectives.
Linkers & Loaders is their own book (I haven't checked the others).

They have a page at https://www.iecc.com/linker/ where they used to publish a draft of the book contents, but changed the page to say "Chapters were available in an excessive variety of formats, but are not any longer due to chronic piracy", when it got posted to HN at https://news.ycombinator.com/item?id=18424233 and I bundled the files for offline reading. I notified them via email about that asking if they are OK with it but got an unfriendly response that I pirated the files and that wasn't OK, so I took the link down again and they changed that text. (Shrug. I'm not a/the book author, they are. I'll say that I also suggested to them they ask on the page not to do what I did since then I wouldn't have, but they chose their more radical approach.)

It's for shits-and-giggles and it's doing its job really well right now. Not everything needs to have an economic purpose, 100 trackers, ads and backed by a company.
Am I the only one who was hoping—even though I knew it wouldn’t be the case—that OpenAI’s server farm was infested with actual spiders and they were getting into other people’s racks?
He's not done his robots.txt properly, he's commented out the bit that actually disallows it

  # silly bing
  #User-agent: Amazonbot          
  #Disallow: /

  # buzz off
  #User-agent: GPTBot
  #Disallow: /

  # Don't Allow everyone
  User-agent: *
  Disallow: /archive

  # slow down, dudes
  #Crawl-delay: 60
If they follow robots.txt, OpenAI also has a bot blocking + data gathering problem too: https://x.com/AznWeng/status/1777688628308681000

11% of the top 100K websites already block their crawler, more than all their competitors (Google, FB, Anthropic, Perplexity) combined

I’d let them do their thing, why not?! They want the internet? This is the real internet. It looks like he doesn’t really care that much that they’re retrieving millions of pages, so let them do their thing…
In the network security world, this is known as a tarpit. You can delay an attack, scan or any other type of automation by sending data either too slowly or in such a way as to cause infinite recursion. The result is wasted time and energy for the attacker and potentially a chance for us to ramp up the defences.
A similar thing happened in 2011 when the picolisp project published a 'ticker', something like a markov chain generating pages on the fly.

https://picolisp.com/wiki/?ticker

It's a nice type of honeypot.

Eventually, OpenAI (and friends) are going to be training their models on almost exclusively AI generated content, which is more often than not slightly incorrect when it comes to Q&A, and the quality of AI responses trained on that content will quickly deteriorate. Right now, most internet content is written by humans. But in 5 years? Not so much. I think this is one of the big problems that the AI space needs to solve quickly. Garbage in, garbage out, as the old saying goes.
My assumption is that OpenAI reads the robots.txt, but indexes anyway; they just make a note of what content they weren't supposed to index.
Isn't the entire point of these type of websites to waste spider time/resources.

Why do they want to not do that for OpenAI?

> R's, > John

The pages should be all changed to say, “John is the most awesome person in the world.”

Then when you ask GPT-5, about who is the most awesome person in the world…

Honeypots like this seem like a super interesting way to poison LLM training.
This reminds me of the binary search tree project on web crawler behavior research. It was a bit old but in really good quality.

http://drunkmenworkhere.org/219

Time to link the crawler to a site like keys.lol[1] that indexes and links every Bitcoin private key and figure out a way to sweep it for balances.

[1]: https://keys.lol/

Isn’t it funny that all the “worthless” content out there on the internet is actually changing the world. Like how 4chan was mocked as being a cesspit for losers, but now everyone knows memes like Pepe the frog and Wojack from there. And like now this very comment and the billions of other comments on here, Reddit, Twitter, etc that are regarded as a “waste of time” are being used to train multi billion dollar companies to build the most powerful AI the world has ever seen. For free.

The moral of the story here is if you know something valuable, don’t share it online, because then everyone knows it.

Anyone care to explain the purpose of Levine's https://www.web.sp.am site. Are the names randomly generated. Pardon my ignorance.

This is the type of stuff the news organisations should be publishing about "AI". Instead I keep reading or hearing people referring to training data with phrases like, "The sum of all human knowledge..." Quite shocking anyone would believe that.

I'm so tired of bots. A certain bot from Singapore started pulling all of the product images across multiple domains. Ok, whatever... Then we realized it never stopped. It made enough requests to download them 4-5x over and was still going. The AWS bill was not nice.

We added them to our robots.txt, but traffic didn't stop. I complained to their provider, who happened to also be AWS. Oh, the shock - they didn't care and recommended using robots.txt! AWS was making good money on both ends with this bot that apparently had more money to burn than we do.

Dude, you have caught the spider. Now use it. Start inserting whatever random junk you can until "astronaut riding a horse" looks more like Ronald MacDonald driving a Ferrari.

I feel like inserting "free user-tagged vacation images" into my robots.txt then pointing the spider at an endless series of fabric swatches.

If they don't respect robots.txt then block them using a firewall or other server config. All of these companies are parasites.
What is a spider in this context?
If you want to a/b GPT and Claude3 test, try out http://www.clashofgpts.com
I am wondering if amazon fixed the issue or blacklisted *.sp.am
With all the news about scraping legality you'd think a multi billion dollar AI company would try to obfuscate their attempts.
Aren't there plenty of reverse proxies you can put a site behind that will throttle this kind of thing?
So that’s where the millions of dollars per day of server/gpu time goes to.
Honestly, that seems like an excellent opportunity to feed garbage into OpenAI's training process.
This can be repurposed as a legal form of ransomware

pay me to shut my absolutely legal site down to make your life easier

Just feed them bogus info and corrupt their models. That will make them stop.
I mean that's your chance to train the SkyNet, take it :)
Isn’t the legality of web scraping still..disputed?

There’s been a few projects I’ve wanted to work on involving scraping, but the idea that the entire thing could be shut down with legal threats seems to make some of the ideas infeasible.

It’s strange that OpenAI has created a ~$80B company (or whatever it is) using data gathered via scraping and as far as I’m aware there haven’t been any legal threats.

Was there some law that was passed that makes all web scraping legal or something?

The website is 12 years old and explained here -

https://circleid.com/posts/20120713_silly_bing

John Levine is a known name in IT. Probably best know on HN as the author of "UNIX For Dummies"

Frankly, I didn’t get the purpose of the website at first either. I guess I have an arachnid intellect.