There is a danger though if certain types of sites are more likely to block GPTBot than others, because that would end up skewing the data set that it trains off, which could have longer term impacts on all the content generated with it. If all the good quality sites block it and the sites full of AI generated junk don't, then that sounds like a downward spiral.
Set up direct Gzip sendfile, and that’ll keep them busy for a while.
If I’m paying for servers by the hour and half of them are run just to keep the bots happy, I’m going to have opinions about who is allowed to visit my pages and for what reason. Mooches don’t have a moral leg to stand on.
I heard many instances of such things.
Wasn't there a case earlier this year that found the Internet Archive liable for copyright infringement?
I'm sure my perspective would be different if I was paying my employees to create unique content for our brand, though.
I'm not sure though. At the end of the day, I think I'd rather information to be free. But that's not a sustainable model in many industries.
AI should not behave as if it has inherent knowledge, there should be aknowledgement similar to, 'i saw on the web the other day'; or 'i was cURLing on coolrecipies.com and it gave me a new idea for halloween cake'
"I saw on the web the other day" is not attribution, it's basically an interjection that makes the conversation smoother. Most of time it means absolutely nothing, except for maybe "it's not my direct experience" at best (and AFAIK LLMs today don't really have any agency, so this disclaimer is moot/noise).
Sure, "I've read an article on Acme Daily" happens, but personally I typically do this as a cue for the listener to cut me off with "ah, yes, I've read this too", saving us both time. Other use case is to give signal about authenticity of the information: not a credit, again - just an indirect indicator of trustworthiness or reliability (when I hear "The Onion reports" - it's surely not about the website, it's about the following being satire). YMMV, of course. I sure want an LLM to write this, but only when it matters to me personally (not the website authors, they aren't a party in our conversation). Just like a human would.
Similarly, I don't think anyone ever said "I've found this on coolrecipes.com" unironically if the conversation is about the recipe instead of a source. Obviously, attributing it to someone both parties know - like a relative, neighbor or a celebrity is a different story - roughly the same as with the news example above. But if you hear "found on coolrecipes.com" it it's most likely an ad. And what I want to say is here is that LLMs today are a breath of fresh air - compared to modern enshittified search engines - specifically because they're not ad- and SEO-ruined (yet). Let us please keep it this way for as long as possible.
ChatGPT is not a human. If a news site was to publish some information on their website without attribution, this would be a problem. ChatGPT is more analogous to a news site then some person you have a conversation with.
the idea is to enable the promptor to review the information, and see for themselves, should they have need.
I'm not being flippant here. We all use massive amounts of reference material to even think the most basic thoughts. Expecting you—or an algorithm—to be able to quote or even know the reference material that let to anything you say is a pretty high bar.
Just like in literature - tropes are standalone entities, and they evolve over time. While a trope can be attributed to some book or author, it evolves, gets mixed up, gets deconstructed and may end up with something barely recognizable. Should LLM be forced to always give a nod to Azimov or Čapek when talking about, let's say, Bender from Futurama - if this is not explicitly relevant to the conversation? I highly doubt so - it would make conversations intolerably stuffy. Talking about virtually any topic would end with a footnotes list multiple pages long.
Our (human) "normal" common sense to attribution is to ditch any and all, unless it's relevant for some reason. Because we generally try to stay focused. People or conversation machines attributing recipes to website addresses is a Black Mirror episode material, a corporate wet dream.
- Why did you block GPTBot?
- Are you aware that your content is scraped, directly copied and otherwise repurposed by other website that don't block GPTBot?
- What are your plans if in future iterations of the GPT model you're going to see that the GPT model has information that you wrote or produced? Are you going to fight it, and if so - how are you going to do that?
I think these are legitimate questions and they are the ones that I would love to hear answers to because I would love nothing more than OpenAI being hamstrung based on the bullshit that they pulled last year with ChatGPT.
Never forget that OpenAI stole the web and has had $11.3B in funding[0] and is seeking another round to place it at a $80-90 billion valuation[1].
[0]: https://www.crunchbase.com/organization/openai/company_finan...
[1]: https://techcrunch.com/2023/09/26/openai-is-reportedly-raisi...
- Yes, and I try to block them as I find them.
- I will fight it if I am capable because yes, OpenAI stole the web, and the more we can make them hurt for doing so, the better.
I block AI scrapers because I don't think that these systems are good for society and I don't want to help make them better.
> Are you aware that your content is scraped, directly copied and otherwise repurposed by other website that don't block GPTBot?
Yes, which is why I've removed my sites from being publicly accessible.
> What are your plans if in future iterations of the GPT model
None. There's nothing I can do about it after it's been ingested, so the only realistic thing to do is to let it go and do my best to prevent it from happening with new stuff.
Although I am not exactly sure about the difference anyway.
Also, if you are going to make an argument, then please do it. I don’t exactly appreciate shallow comments when I have plenty of evidence to show as a rebuttal.
As they are at the moment, OpenAI are parasites.
A human reads something for free or "free" (if they're a product for advertisers or data collectors). And no one bats an eye if they retell it to others. There are careers of this - teachers, instructors, advisors.
But if it's a a machine (or a superhuman - I want to hope hope we'll see AGI someday) that can do this at scale, the whole world is suddenly upside-down. I honestly have no idea how things should be (I have opinions, but I cannot really validate them, so they're just... opinions), but I find it interesting.
Being a pessimist and strong believer than humanity as a whole leans towards shittiest-but-cheapest solutions (down to a certain threshold, of course, but the plank is rarely high) my guess is that those websites will eventually try to monetize LLMs the way they do this with humans - by injecting ads in the information provided. Fear the day new breed of SEOs will start spamming, inventing fancier and fancier techniques of - ahem - enriching the models with customer-oriented targeted promotional materials (or whatever). And so the history will repeat itself.
I guess another way I see it is like the site has put a banner at the top saying I can read their content, but can't use anything I learn from it, and can't tell you or anyone else about anything I've read. I get that in this case the 'me' is an algorithm and not a person, but does that really matter? If I read a lot of info on welding and then offer a 'learn to weld' class, how much does that differ from what the openai algorithm is doing. (And if it is public, are these same sites going to welcome the bots? I doubt it)
It also seems rather reactive and not well thought out. To include a quote from an Ars article (which gptbot can't read):
> As a thought experiment, imagine an online business declaring that it didn't want its website indexed by Google in the year 2002—a self-defeating move when that was the most popular on-ramp for finding information online.[1]
Or a couple quote Stephen King on another site that gptbot can't read[2]:
> I have said in one of my few forays into nonfiction (On Writing) that you can’t learn to write unless you’re a reader, and unless you read a lot.
> Would I forbid the teaching (if that is the word) of my stories to computers? Not even if I could. I might as well be King Canute, forbidding the tide to come in. Or a Luddite trying to stop industrial progress by hammering a steam loom to pieces.
As I said, to each their own. I just would have expected a site like NPR, that lists it's mission as "to create a more informed public", to not take actions that work directly against that goal.
[1] https://arstechnica.com/information-technology/2023/08/opena...
[2] https://www.theatlantic.com/books/archive/2023/08/stephen-ki...
Thank you for explaining.
More: https://tomaszs2.medium.com/ai-may-pirate-music-and-movies-1...