back
985 comments
I see symptoms of this all the time. For example, it's a weekly annoyance for folks to pop into /r/strava to showcase their vibe-coded app that uses the Strava API to do $THING. Then someone invariably points out that an existing app (or even Strava itself) already does $THING, and often it's free. I don't mean to be negative, I think it's great that people are building useful niche software and I don't blame them for wanting to share it. A significant part of the problem is that it's much harder nowadays to find "prior art" because keyword/boolean web searches have been FUBAR.
I've seen this across so many other subreddits as well /r/formula1 /r/cycling etc

And its always very obvious that its an AI developed front end, they all look so similar.

It's certainly software being developed at speed by anyone with a level of tech interest. Which is great but its frustrating when the native solution provides said feature.

Hopefully now that AI provides quick answers, web search can go back to providing and respecting boolean and more advanced techniques. If web search is no longer the go to tool for most users, then search can do better what it can do uniquely?
Same phenomenon for the game Satisfactory, lately I've seen so many vibe-coded tools that just do the exact same thing as already existing tools.
I've used a lot of free features that I would have implemented differently, and now I can (if I weren't so lazy).
> A significant part of the problem is that it's much harder nowadays to find "prior art" because keyword/boolean web searches have been FUBAR.

I’m surprised this isn’t discussed more. I find most major search engines are basically unusable since the top search results are overwhelmed with AI slop. Trying to find anything useful is like looking for a key in the mud.

My sister, a journalist, mentioned to me that she only uses google search because she had learned how to get information typically only Google indexed in the country she lives in, in a way it was not exposed on chat bots. She often has to search for information like Old govt forms released as public record with a fixed a certain format photo scanned into a pdf and indexed by Google were often on the second page of the search and beyond. But they are there. She knew how the forms looked and what bigrans and trigrams matching a certain part of form for a certain piece of information to search for and Google search has it. Like an official order on a tender notice for some government department which is no longer in the .gov.* website gave her the official's name and then she could track down who to contact in an office... ChatGPT and other bots don't have it. Some how all these government documents became part of the government record and are the key for her to do her job.

I sincerely hope google wont stop indexing that stuff just because of a PM in search "de/re-prioritizing" ranking in a way that makes this impossible.

This is a conundrum I always found interesting. If you know how to use a search engine (i.e. knowing how to use operators and structure a search query), you're almost always able to find what you're looking for very quickly, and in most cases (well, before SEO), the results are high-quality. You'll spend the same amount of time trying to fact-check an LLM (since you'll likely skim the articles it used in generating its response _which you would have done anyway if you used the search engine directly_).

I actually took a (required) class in middle school that taught us how to use a library. Amongst other things, the librarian taught us how to use Google effectively. Everything I learned then (this was in the early 2000s) still works today, since the process of using a search engine hasn't changed very much since its inception.

So many people never learned (or never cared about learning) how to use a search engine, thus why we're here today.

Would be interesting to know if those same long tail results come up in Alt-Power [0], which also uses Google's index. So far I get what I ask for, but would be reassuring to know the whole long tail index is indeed shared and I'm not missing relevant results.

[0]: https://altpower.app

> just because of a PM in search "de/re-prioritizing" ranking in a way that makes this impossible

I could be wrong, but I'm more inclined to say these directives are coming from leadership/advertising dollars rather than a seemingly rogue PM

After publishers successfully sued the Internet Archive over its digital lending program, calling it unauthorized copying

No. The court specifically determined that the Internet Archive was guilty of unauthorized copying. It was not simply an unfounded or unproven allegation. The Authors Guild, the National Writers Union, the European Writers Council, and the Society of Authors in the UK all came out against the Internet Archive, and supported the suit.

Each new restriction limits the archive’s ability to act as a comprehensive backstop.

This self-inflicted damage to the wayback machine is the real tragedy of this entire affair. When IA was asked to stop CDL - many times - founder Brewster Kahle continued. The National Writers Union tried to open a dialogue as early as 2010 but was ignored:

The Internet Archive says it would rather talk with writers individually than talk to the NWU or other writers’ organizations. But requests by NWU members to talk to or meet with the Internet Archive have been ignored or rebuffed.

https://nwu.org/nwu-denounces-cdl/

When the requests to abandon CDL turned into demands, Kahle dug in his heels. When the inevitable lawsuits followed, and IA lost, he insisted that he was still in the right and plowed ahead with appeals. And here we are today.

> No. The court specifically determined that the Internet Archive was guilty of unauthorized copying.

You're not wrong, but you're treating “guilty of unauthorized copying” as a statement of physical fact when in reality it just means it falls under an arbitrary rule invented by humans (namely, the law that defines unauthorized copying). This rule is ambiguous at its edges because it's not written as an algorithm or equation. It was perfectly reasonable for Kahle to believe that the rule can be interpreted in a way that it wouldn't apply and, by dragging it through the courts, have that interpretation be made the established one.

Even though the court has now established a competing interpretation, it is still not unreasonable to ask whether the law is fair and just under this interpretation. I feel that it isn't and should be changed.

I spent the last three days (off and on) using Gemini to configure my edge router 4 with my iOS devices on a vpn and it's been awesome. In the past I'd do a google search and read a few sources of documentation, do another google search and read another set of documentation. Now, Gemini aggregates multiple pages together so all of the work of reading source docs from multiple locations is now n a single step.

Oh, I should mention though. There was no advertising at all. They didn't make any money off me. It was 100% Gemini which I recognize as not long-term feasible.

I called this maybe 3y ago, but I think so did everyone else that was sane. Sure, we get immense value from AI, but indiscriminately injecting into everything, the one thing we know to be unreliable above the threshold we used to fire people for, is probably the greatest undoing of all the good companies like Google brought to the internet. I mean what a way to destroy your legacy of democratizing information. The amount of harm (direct and indirect) this will cause, and the cost to return to baseline will be so immense, and yet we will not be able to point to the root cause. They won't be there to take responsibility.
Funny, I was just thinking this morning that Google searches are absolutely horrible these days. It's like it has amnesia, a lot of recent history seems to be just gone. Especially on non US specific sites too.
I occasionally use Google Search when DuckDuckGo fails to give me relevant. Almost always, Google has better results.

Though I can find its AI answers annoying aggressive. I'll look up like two search terms and the AI will bullshit multiple paragraphs out of despite having zero context of what I am looking for.

DuckDuckGo seems to have detection of whether it should give an AI answer. And it allows you to have more granular control of when you want to get an AI answer. And is overall less distracting than Google's.

The article touches on something that I've been thinking about with regards to Google's AI strategy; the automatically-generated AI search summaries are not great. They very frequently confidently misinterpret what the user is searching for and generate half a page of useless information that pushes actual results down the page, and they are occasionally hilariously incorrect, with hallucinated facts.

This is probably a difficult-to-solve problem; given that they generate billions of these a day, not even Google can afford to devote enough compute to each query to reliably generate quality results. You can see this by selecting the "AI mode" from the search interface after getting the mediocre summary - the results are much better and generally perfectly usable. Though even that is probably a special minimal-compute version of the lowest tier of Gemini, it's still maybe an order of magnitude more capable than whatever generates the search summaries.

The bigger problem is that these search summaries are the default and by far the most common interaction that the general public has with "AI", and because this experience sucks, they just assume that all LLMs are similarly stupid and mostly useless. In non-technical spaces I frequently see the argument that "AI" is not useful for anything, all it generates is garbage hallucinations, and almost invariably they cite some actual terrible experience with the Google AI search summary. I would argue that the strategy of adding LLM summaries to every search is the worst of both worlds - it makes classic search worse while poisoning users against the idea of actual LLM-assisted search.

I feel like collecting, curating, and protecting high quality corpuses of "truth" is going to become increasingly important for high quality AI.

There will come a day (and probably soon) when "training on the public internet" (Reddit, etc) will taint your model with metric tons of corporate contamination, political poison, and other adversarial content intentionally crafted to bias AIs for various reasons (corporate gain, geopolitical information warfare, etc). Basically the AI-equivalent of SEO.

Gemini has been a hilarious companion to my while I fixed the balance shaft chain guides in my old Mitsubishi triton (mighty max for US readers).

First it told me I could just remove said balance shaft chain as an emergency repair. Sorry Gemini, it also drives the oil pump.

Then it told me I could remove the water contaminated oil caused by removing the timing case by filling the crankcase with hot, soapy water and running the engine. Lord no.

Then it gave the wrong instructions for putting new gears on the balance shafts which meant the chain guides didn’t align with the chain. I’ll do it my way thanks Gemini.

The rest of the mistakes are too trivial to recount and sure it’s a pretty obscure subject but if I trusted it with a topic I’m not familiar with there is a huge potential for damage if you blindly follow it’s overconfidence. I miss normal searching.

Kagi search today is better than Google search ever was.

And it’s clear that Google’s Ad model ultimately created a priority inversion. The advertisers became the customer.

I am so glad Kagi came along with a business model that is actually working.

I don't think it's really an "AI problem", we just got to the "worse" part of the "Worse is better".

Back in the day one of competitors of the World Wide Web was Project Xanadu. Project Xanadu was supposed to address the concerns like content persistence and version management within the core design. As such it was much more complex, opinionated and centralized.

WWW on the other hand comes with no guarantees - you might get a document in response to a HTTP request, and that's it. But WWW service can be rolled out in a completely permissionless way, and is quite simple - effectively, the contents of the file system can be shared with the world, so e.g. a document can be published just by putting its file into a particular directory within the file system.

Thus Web could get to a "good enough" state much faster and quickly spread all over the world. But its permissionlessness and simplicity lead to downsides: impersistence and chaos of broken links, web search provided by mega-corporations, etc.

WWW evolution was, unfortunately, not "incentive compatible" with features like advanced persistence and identification clarity: there was much more focus on entertainment content and ads

> While the web has always been organized around intermediaries that shape what survives online and who sees it,

This statement, from the sixth paragraph of the article, is something that I would have liked to see addressed more in the article. The article implies that this is something that must always be true, or cannot be changed, and simply focuses on how we could have better/better funded/better protected intermediaries (AKA gatekeepers), and doesn't discuss the possibility of an internet (or part of the internet) without gatekeepers (and doesn't ask if it has ever existed/does exist/should exist)

AI has killed reading-anything-written-after-AI for me. Due to this effect it is probably the worst invention in human history or pre-history.
I'm working on using the OpenZIM format to archive the web and to make the wikis seedable (and locally hostable for LLMs) so that the ongoing cat and mouse game anubis defense can stop.

My hope is that with the torrent protocol we can make the archived knowledge discoverable and seedable, because currently there's only the web archive and the kiwix download servers for archived contents. Both of them still are centralized servers that bear the cost of hosting those files.

- [1] https://github.com/cookiengineer/gozim

- [2] https://github.com/cookiengineer/zimdex

AI will kill the internet because it is killing the incentive to make it. It is an industrial-strength example of why we don’t allow stealing.
Google wasn't great for a very long time. Switched to Duckduckgo years ago. I just love the bangs, because I tend to go to sources I trust anyway.

Doesn't mean I'm not also using duck.ai. It makes searching faster and more targeted. But then it's still giving me links to verify and is actually more limited, which means less hallucination and more directly going to the sources. Also it avoids having to open five pages first which all either sell your data or want you to pay.

I don't see the web or the internet dying yet. Just a lot of people not using the right tools and having a harder time accessing what's useful. But that hasn't started with AI.

I was asking Claude yesterday about some specific roman history when it appeared to hallucinate a fact I knew not to be true - when I questioned it, it said it sourced it from an Encyclopedia Brittanica page that was "flagged" as being an AI generated summary of their actual content, and admitted that the fact I challenged was not historically supported.

So I guess we have entered the age of AI-generated "alternate facts" - one AI citing another AI's hallucinations as fact.

we're building the world's largest library and then locking the doors, letting the bots photocopy everything before the lights go out.
From a user perspective, Google search is the most useful it has been in years, though that doesn't feel entirely like intentional improvement, just a lucky side effect of the move to "AI mode".

And yes, if you take what the AI tells you at face value it could be wrong. But if you are aware of this and aware of the ways in which LLMs are likely to shit the bed, it is quicker to get from request to useful information than it has been with Google search since like 2017.

And also, yes, the old balance of Google driving clicks to sites that will then generate revenue off more Google Ads being shown after you click through to them creating a virtuous cycle is completely busted, and that sucks. It does not impact me directly but it certainly seems like unless a better system is devised that it is one of a few ways in which AI is likely to stall out its own training funnel.

Overbroad claim. Dramatic corollary

This clickbaity headline format cannot die fast enough

The web was getting kind of useless before AI crashed the party. This is why now curated content is key - newsletters, for example, is how I find most of my content.
I think it's a chance to return how the old web was, as the human web return to a small underdog and get splitted from the AI web.
Google switching to hallucinating AI summaries has been to me an absolutely shocking abdication of care for both their users and their own reputation.

It extends beyond search as well. I have had multiple incorrect Gmail summaries that, if I had only read them instead of the actual email, would have resulted in financial harm.

And the same holds for documentation about anything. Documentation is gone. Dead. Reference documentation, output from Doxygen or similar, and many other things that could be searched for hints on how to implement or generally do stuff. It's gone. When writing a simple Python script today and wondering about how an API for some library works, I ask AI, because there is no (findable) documentation anymore. (And yes, I usually still like to write it myself, but it's basically no difference: the AI writing that Python script would be the same point: docs are dead and gone.)

This is really scary and it is progressing fast.

Really well written article and interesting. I'm not sure how I feel about a governmental policy over retaining access to information though, the information is provides by us and we pay for the infrastructure, the idea that there must be some form of retention policy makes me feel uneasy and doesn't really fit with the analogy of governments maintaining roads.

I'm on both sides. I hate dead links but I'd hate a policy that made me responsible for them without any compensation in the first place. It would probably make me stop producing at all

This source alone - Multiple ads. Overlays. Subscribe nags. What appear to be more clickbait ads. (Yes, I know Adblock exists. That isn't the point.)

The way the web is these days, I think I'm actually ok with AI eating it.

Using an AI nowadays reminds me of the old Gopher days - you get simple, plain text. Perhaps I'm just old enough to miss that.

We already know how to compute sunset times accurately and doing so requires one millionth the computer power of LLM inference. It could even be cached for most big cities for every day of the year.

https://www.timeanddate.com/sun/

The mystery is why Google doesn't just route such requests to the known algorithm. It would be a lot cheaper for them and it wouldn't risk reputational damage.

More than 25 years ago, the Web wiped out a lot of the things I loved. May it be devoured! But I have my doubts: there’s probably not much left of it anyway.
Perhaps we should let the content economy crash so that the big monsters eating it starve and die. A new content economy could be built on their carcasses instead.
The concept of an almanac seems relevant again: a yearly printed book with verified, accurate information. No manipulation at a later date, no AI hallucinations, etc.

The most famous one was probably Benjamin Franklin’s:

https://en.wikipedia.org/wiki/Poor_Richard%27s_Almanack

Paper encyclopedias might make a comeback for the same reason.

It's hard to get past the beginning of this article and take it at all seriously. The quote someone who missed a sunset because they asked google and supposedly got the wrong time... but they didn't think to look at the sun or lack thereof to check? Also when I put "when does the sun set today" I get a single exact figure at the top of my results, not from AI, which is honestly the best kind of result – an exact correct answer.
> Wikipedia has become the infrastructure of its own demise: dwindling traffic means attention and donations no longer reliably flow back to the encyclopedia to keep it alive.

It’s too bad this will drown out the feedback from folks like me, who ended their longtime recurring donation because of their resistance to their employees unionizing.

My father used to print websites up and put them in a 3-ring binder. Now it seems more prescient than anachronistic.
Gemini's highest tier of plan has the utility of Google search circa 2010. I'm paying $400 for the privilege.
What should be concerning to Westernized nations is that as decent stewards like Google fails; the information pipeline into our culture and minds still remain.

i.e., it's not just the collective memory going away, but what will easily replace it and who will be motivated to influence.

I find Brave search is superior to Google, particularly in linking me to more useful references.
So let it be. Why don't we build another place for our memories to go? Why don't we build the _unstructured_ internet, where the intelligence is not in the mind but in the eye and the pleasure is in finding not in disseminating?
I see only one simple solution (though we should discuss the more complex ones): if Google directly answers a search query, then it must be held accountable for it, for better or for worse, and therefore assume all the benefits and (legal) liabilities that this entails
The internet might not be dead, but large parts of it are gangrenous and necrotic