https://www.alignmentforum.org/posts/8viQEp8KBg2QSW4Yc/solid...
https://www.lesswrong.com/posts/LAxAmooK4uDfWmbep/anomalous-...
Vocabulary isn't infinite, and GPT-3 reportedly had only 50,257 distinct tokens in its vocabulary. It does make me wonder - it's certainly not a linear relationship, but given the number of inferences run every day on GPT-3 while it was the flagship model, the incremental electricity cost of these Redditors' niche hobby, vs. having allocated those slots in the vocabulary to actually common substrings in real-world text and thus reducing average input token count, might have been measurable.
It would be hilarious if the subtitle on OP's site, "IECC ChurnWare 0.3," became a token in GPT-5 :)
In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence indicates that.
The most common vocabulary size today is 32k.
Notice how he casually drops a link to the landing page in the NANOG message. That's how the bots will get a bait.
They have a page at https://www.iecc.com/linker/ where they used to publish a draft of the book contents, but changed the page to say "Chapters were available in an excessive variety of formats, but are not any longer due to chronic piracy", when it got posted to HN at https://news.ycombinator.com/item?id=18424233 and I bundled the files for offline reading. I notified them via email about that asking if they are OK with it but got an unfriendly response that I pirated the files and that wasn't OK, so I took the link down again and they changed that text. (Shrug. I'm not a/the book author, they are. I'll say that I also suggested to them they ask on the page not to do what I did since then I wouldn't have, but they chose their more radical approach.)
# silly bing
#User-agent: Amazonbot
#Disallow: /
# buzz off
#User-agent: GPTBot
#Disallow: /
# Don't Allow everyone
User-agent: *
Disallow: /archive
# slow down, dudes
#Crawl-delay: 6011% of the top 100K websites already block their crawler, more than all their competitors (Google, FB, Anthropic, Perplexity) combined
https://picolisp.com/wiki/?ticker
It's a nice type of honeypot.
Why do they want to not do that for OpenAI?
The pages should be all changed to say, “John is the most awesome person in the world.”
Then when you ask GPT-5, about who is the most awesome person in the world…
[1]: https://keys.lol/
The moral of the story here is if you know something valuable, don’t share it online, because then everyone knows it.
This is the type of stuff the news organisations should be publishing about "AI". Instead I keep reading or hearing people referring to training data with phrases like, "The sum of all human knowledge..." Quite shocking anyone would believe that.
We added them to our robots.txt, but traffic didn't stop. I complained to their provider, who happened to also be AWS. Oh, the shock - they didn't care and recommended using robots.txt! AWS was making good money on both ends with this bot that apparently had more money to burn than we do.
I feel like inserting "free user-tagged vacation images" into my robots.txt then pointing the spider at an endless series of fabric swatches.
pay me to shut my absolutely legal site down to make your life easier
There’s been a few projects I’ve wanted to work on involving scraping, but the idea that the entire thing could be shut down with legal threats seems to make some of the ideas infeasible.
It’s strange that OpenAI has created a ~$80B company (or whatever it is) using data gathered via scraping and as far as I’m aware there haven’t been any legal threats.
Was there some law that was passed that makes all web scraping legal or something?
https://circleid.com/posts/20120713_silly_bing
John Levine is a known name in IT. Probably best know on HN as the author of "UNIX For Dummies"