back
377 comments
So spammers have latched onto your search engine because they are getting useful results. They are able to systematically discover websites built on certain platforms that allow users to post content containing links, which they can target for link spam. It is very difficult to fight this on a technical level because there is an entire industry built around blackhat SEO, with all kinds of softwares and services dedicated to thwarting your defensive efforts. Even Google struggles to keep up with this.

However, they are also systematically feeding you their footprint lists. I imagine you could put together a footprint blacklist pretty quickly, and just stop returning results for any obvious spam queries like those containing "powered by wordpress".

It's not a very elegant solution I'll admit. It won't stop the bots from trying, and you may have to circle back periodically to add new footprints as they surface. But it's a potentially quick and easy way to stop rewarding their efforts, and the blackhat world is pretty used to burning out their resources so hopefully they will figure out it's a dead end and move on.

> So spammers have latched onto your search engine because they are getting useful results.

I'm not sure about this. At least with my search engine, it doesn't really seem to matter what response they get, I don't even think they look at the responses. They keep hammering away with tens of thousands of queries per day with the requests even though they've seen nothing but HTTP Status 403 since last October or so.

My best guess is they're going after search engines in general in case they forward queries to google, in order to manipulate their typeahead suggestions.

Or put those query results behind an anti-bot/"capcha" test.
Considering that as of Mar 12, this search engine only has 1001 sites indexed, I am not sure how useful this site is for getting SEO backlinks. Speaking of which, are backlinks still a thing these days?
If the confidence was high enough, perhaps return garbage data?
> It is very difficult to fight this on a technical level

It is when your base assumption is that you won't hire outside of engineering. There are more bored teenagers with phones than people creating quality content, so I'm not sure why you wouldn't just brute force checks against bad actors.

just to throw out ideas: What if he decided to charge for each search?, say 1 cent or so. Users could purchase them in bulk, say 100 searches for a 1$.

The world is getting more and more desperate for a better search engine. the day may come, when people are willing to pay for better results.

what is the end goal here? i understand it's about making money somewhere down the road. but how?
Since everyone in this thread wants to jump down OP's throat about the quality of his web site, another interesting search engine is millionshort.com, which allows you to filter out the top N web sites from the results of your search. It's a great tool for looking past sites with good SEO; all you have to do is fiddle with the value of N.

For example, searching for "electronic music box" as /u/ajnin suggested, with the top 100K web sites removed from the results, filters out the following:

> These 23 sites were removed from your results:

> alibaba.com (1 result removed)

> aliexpress.com (1 result removed)

> allaboutcircuits.com (1 result removed)

> amazon.com (2 result removed)

> apple.com (1 result removed)

> bestreviews.com (1 result removed)

> ebay.com (1 result removed)

> etsy.com (2 result removed)

> facebook.com (1 result removed)

> instructables.com (2 result removed)

> lightinthebox.com (2 result removed)

> lumberjocks.com (1 result removed)

> mapquest.com (1 result removed)

> reverb.com (1 result removed)

> twitter.com (1 result removed)

> wikipedia.org (1 result removed)

> yelp.com (1 result removed)

> youtube.com (2 result removed)

And the top result ends up being https://midiguy.com/.

That's an outstanding concept. One problem though: wouldn't it also filter out high quality curated results?
Million Short also has an option to remove only e-commerce results which is invaluable if you still want results from sites like Twitter, Wikipedia and YouTube but don't want online shopping spam.
This made me curious to try that search engine so I typed "electronic music box" (first thing that came to mind). As far as I can tell none or the 10+ pages of results include all those 3 words. I mean, you might not have any relevant sites in your database (likely if there are only 1000 sites or so as another of your blog posts imply), and I understand you want to show some result to the user, but if I want irrelevant links I might as well go to google.com...
You mention the "Dead Internet Theory" (not heard that phrase before!).

I agree: the WWW Internet is dead, that is your problem. No-one visits websites anymore, everyone has moved to the 10 biggest websites and all data is now siloed there.

If I want to search for something topical and relevant, I go to Facebook, Twitter, Reddit, HackerNews, Instagram, Google Maps, Discord etc.

The general Internet is dead: it's just legacy content and spam.

If you think it's bad for you, imagine what it is like for Google Search! Their entire business is indexing a medium which no longer has any relevancy. People complain that Google no longer delivers good results. But what can Google do? The "good content" is no longer available for them to index.

Want to become rich? Make a search engine which indexes the fresh relevant data from the big siloed websites, and ignores the general dead Internet.

Ona tangential note, I remember a time when Google had the option to search only for 'discussions'. The results were amazing and accurate as it scoured online forums. Almost all issue I had (was following the rooting scene closely back then) were quickly resolved. Then suddenly it got removed for reasons unknown to me. Anyone knows if it's replicatable today?
> I didn’t notice at first because the web analytics only shows real users, and the unusual activity could only be seen by looking at the server logs.

Sounds like everyone blocking analytics (Plausible in this case), e.g. myself just now, is lumped in with spam bots.

Of course, analytics blocking can’t meaningfully swing the ~99.99% statistic.

I'm disappointed that Search My Site isn't seeing many legitimate viewers.

Just wanted you to know that I'm a fan. I love reading peoples personal websites, and Search My Site has been great for discoverability. I visit the Newest Pages and Browse Sites pages once or twice a week to check out the new sites being indexed.

I don't know what the answer is to the spam bots, but you do have some real visitors out there. :)

This guy throws multiple reasons/conspiracies out there on why the website is really struggling to gain literally any sort of traction. Web is all bots, search engines not promoting competitors and being drowned out by SEO spam, yet he's failing to see the most obvious reason... the reason nearly all websites don't gain traction...

Because it's a bad website. It provides no value to the user. I put in a few search terms and had no relevant search results back. What use is a search engine that can't find what I'm searching for?

Maybe if that was improved he may see traction.

Search traffic has always been mostly automated spam bots.

Even back in the Open Directory Days when we powered part of search.netscape.com I estimated 80+% of all search traffic was automated. At least most of it self-identified with the same Java useragent.

Later when working Topix, despite being a news search engine, most traffic was bot traffic. Most included the word “mortgage” in the query. Topix specialized in localized content, and that was very popular for SEO scrapers.

Lastly at Blekko, I estimate 90+% of traffic was automated. By then maybe half or more learned to change the user agent. Most used HTTP/1.0, a dead giveaway as no browser still uses 1.0. This was a major aspect in Blekko's load shedding strategy. If the servers started to get overloaded, we'd start bouncing suspected bot traffic to a redirect that would show in the logs. If there was a human with a modern browser running javascript on the other end, would get redirect to a link that wouldn't get bounced. I would check the logs weekly to see if any humans got caught. None ever did. This was a huge monetary savings, you only need 1/10th the servers if you can safely ignore the bots.

Often it's endless repetition of the same keywords in a random order with a place name appended, or prepended, or inserted. over and over. Often variations on known monetizatable SEO keywords. However, much of it doesn't make any sense.

I don't have any insight into Google's numbers but I would conservatively estimate 95% or more of all their queries are automated bots and not humans. And the level of spy-vs-spy going on for Google CPU resources vs SEO bots is probably pretty evolved by now. I stopped tracking many years ago when Google switched to densely packed obfuscated javascript for page renders. Maybe this is part of why automated queries are so high across the web, maybe google is too hard to crack for most.

One day we'll have an internet for humans exclusively. On another note, with 160K requests / day from bots you could of course simply block the bots structurally assuming they are nice enough to identify themselves. Block all of AWS and Google, Russia, China, NK and a couple of other bot hot spots and the service may well become more successful for regular users because they get faster results. Bots can afford to wait, humans are often impatient. And with 2 hits / second by bots that may well become a factor.
This is for comment spam.

It's trying to find a long tail of popular but not top listed blogs for the purpose of posting comments with the much desired links to the SEO target.

If the internet is dead, is there anything left that's "alive"? The mobile app stores are also filled with crap[0] and it seems that the ratio of spam content vs real content is getting close to infinity.

[0]: https://youtu.be/E8Lhqri8tZk - 1,500 Slot Machines Walk into a Bar: Adventures in Quantity Over Quality

I have had very interesting conversations with people who are "casual" users of internet. They are still finding the results of the likes of Google, bing and duckduckgo perfectly suitable. Maybe it's most of us here who have different needs to what's available.
Mojeek member here. We have always had a high level of spam bots; as any search engine/service will have. It's a constant battle to fend off new bots; folks can always use try out our API rather than freeloading, and some do. Many obviously do not. We are taking a look at whether things have also changed for us since mid-April 2022.
That's really interesting... and sad. For what it's worth, I've noticed comment bots dramatically increase over the last year too. They have always been there, but looking at Reddit, YouTube, etc, now there seem to be 10x more than there were a few years earlier. Even on HN it has gotten worse.
i recently built a habitat for spam bots, they eventually found it and now post peacefully

https://upstairs.treehouse.telnet.asia/pharm/cylohexapine

I had to put my search engine behind Cloudflare to deal with this. Like the volume grew to about 10x the traffic I saw sitting at the front page of Hacker News for a full week.
If this was the spam for a search engine (almost) nobody uses, it makes you wonder how much abuse the major search engines face
My key takeaways:

1. Almost all searches on my independent search engine are now from SEO spam bots

2. In summary, if they break through the current reverse proxy level protection, options include an invisible ReCAPTCHA (but given I’ve sometimes 160,000 requests a day I’d be well over the 1,000,000 a month free tier limit), requiring JavaScript as per the web analytics or some Cross Site Request Forgery style protection (but those would place much more load on the servers), or CloudFlare (but the searchmysite.net spider is still currently blocked by CloudFlare as per Some of the challenges of building an internet search)

3. If you were into conspiracy theories you could claim that the major search engines were trying to stifle the competition, but a more realistic explanation is simply that searchmysite.net is being drowned out by SEO spam

4. If I’d had a decent amount of real users visiting and never returning I could reasonably conclude that updating the blog wasn’t the most productive use of my time and effort, but without any real users in the first place it is hard to gauge whether people like it or not

My own independent search engine, https://www.locserendipity.com, is seeing similar trends.

Well, the first two links loaded for a search for "magic the gathering" are 404s. The "Random" link at the bottom 403s. The search engine feels broken.
I run a data aggregation company that has a fairly advanced scraping infrastructure for collecting data across the web. Having built the scraping side, I'm pretty familiar with most of the strategies for avoiding bot detection.

Coming from that perspective, detecting and stopping at least the majority of bots out there is fairly doable, and I put together a rudimentary thing for a side project.

The core of it uses an IP API for looking up the requesting IP to identify the country and if it's coming from a data center, VPN, Tor, etc. If it passes that, I trigger Google Captcha to show up. Lastly, I track IPs that make it through and have some basic rules in place to try to detect patterns and block offenders that way.

There's a bunch more stuff you can check for, but the core of it is basically filtering out data center traffic to minimize the requests going to Google Captcha.

Complete SEO noob here. Can someone help explain what these bots are trying to achieve? There is mention in the blog that they're trying to uncover ad free content.
IMHO what you should try is excluding all sites with excessive third-party cookies, sluggish performance, and too many ads. That will slice the index down by 80% probably but it would be a really nice thing to see. It might push out low quality SEO results for a couple of years.
In late April up to now, Wiby (a small mostly unheard of search engine) began having the exact same issue. Tens of thousands of the exact same type of "powered by..." requests coming from thousands of IPs. They are using a tool called QHub.
Spammers badly need spam-free content so they can mix some legitimate links with the junk they spew.

One great Black Hat SEO trick is to find where your competitors are getting clean links and insert your own links there so they do your spamming for you.

Random tangent related to SEO I am so annoyed with random overseas companies faking local businesses by SEO

They will get your call and you can tell where you're talking to immediately. Then they get a quote that's increased to get their cut... That's then carried out by an actual local company. It's annoying because the websites at the top of the search appear local.

A specific example is when you're looking for a towing company.

Type in your state/city towing company, more than likely the top results/websites pinned to GMaps are not local-based.

I think many people in the comments here, and most users, are missing that you index a SMALL subset of the web. This leads to people running a default test search, finding no results, and concluding your search engine is bad, and leaving.

While you imply that in the search page, obviously it's not clear enough.

Maybe add "this search engine only searches a small set of user submitted sites. Click <here> for the list. Or <here> to add your site."

>I noted that there had been multiple weeks where not one single real person had visited a single blog entry for the whole week

The site is not on https://searchengine.party/ nor on seirdy.one's overview. Apart from the blog, how could users find that engine?

Is there some place where new search engines are announced and where new search engines band together to make themselves heard?

I created a temporary email service that was being used by about 10k users / week. Then several weeks ago, the number of users started growing like crazy up to about 60k users a day. Then we checked the recent email activity and 60k / 65k emails were from a social networking site.

Seems our service was being used to create fake bot accounts. The newly created accounts were obvious fakes. Rather than deal with the issue, we just shut the service off.

This is an awsome website that I was not aware of!
I've found an effective way to filter bots is to check their headers. The order and capitalization of the headers. Browsers will always send headers over with the same capitalization and order (unless some plugin interfering is installed). Check your bot traffic and see what browser it's masquerading as, then compare some headers that are sent over by that browser and version. For example 'Content-type' is coming over as 'content-type' or 'Content-Type'. Or the bot is not sending the Accepts header whereas Chrome always sends it. A few small checks on the headers can help easily identify a lot of the bots.

Chances are they're not going to take the time to figure out your specific site's mitigation measures and just let the bot fail crawling or move on.

It is my experience that SEO bots are increasingly ignoring robots.txt entries disallowing them from crawling our sites. Last week we noticed several doing this. I don't mind naming names - semrush, something called grapeshot crawler, something else called blex bot, and moz dotbot. Anyone else having the same experience?
People will only use your product if they know about it and perceive value in it. How do people know about it, and why would they want to use it?

On "Most of the tiny number of real users have come from links posted to places like Hacker News, and there is almost no organic traffic from other search engines" - Organic traffic comes from word of mouth. Are people talking about your site? If they're not, you're not gonna see organic traffic. You could do what others do and pay some influencers to advertise your site, but that's expensive and not as scalable as "real" buzz. Is your product exciting or controversial? If not, why would people talk about it?

Your homepage's tag line is "Open source search engine and search as a service for personal and independent websites." A regular person's eyes would glaze as they try to figure out what this means. Given some time they might put together the words "search engine" and "personal" and "websites" and figure this is a blog search engine. So just say that.

The "Newest Pages" section is a fun novelty, but after a few minutes the novelty wears off.

The "Browse Sites" section is almost useful. Next to the list of sites I see some tags. Why isn't a heatmap of the tags the first thing I see? That would be way more useful than a paginated list of random sites.

Your "About" page lists "community-based approach to content curation". This is the most exciting aspect of the whole endeavor, so add that to your front page blurb ("Community search engine"). You would probably do well to build a real community around it, for example with a forum or chat system (GitHub Discussions does not count). A SubReddit would be an easy way to bootstrap this and later move it to your own hosted forum.

You'll probably need a very complicated moderation system if this thing takes off.

To me it looks like some popular spamming software (like thebestspinner, etc) just integrated you and now everyone who is the software is now hitting your site.

The good news in this case is that's it'll be easy to spot the pattern and block it, the bad news is you're entering a never-ending cat and mouse game.

If there's a game to play, people will write software to play it for their profit.

I guess it's back to web rings.

Only solution is a webring based federated search engine.

1. You just put /webring.txt to your website. It shows links to other websites with a hard limit of 100 websites.

2. To combat spam and bots, search engine does accept blocklist as an input. So other people can curate the content.

3. People can personally rank the websites they like, so webring of the said website gets ranked higher for that specific user. This can be a community effort too.

4. Search engine itself should be under a commercial license so that other people can keep building it and add ads if they want to commercialise it.

I’m too busy to spend time with this but perhaps one day I can start coding it.

I’m convinced that search engine model of early internet is just dead, webrings are the way forward.

I am a total ignorant about search engines and I have a question after seeing all types of comments and projects popping up lately and criticizing Google results which is if it is realistic to think that something similar to Google could exist.

It seems to me that there are all sort of tools out there to do so such as all the public NLP implementations, vector search engines, ... and I wonder if it is that not everything that is needed is truly available, it is a matter of the needed resources to have something working or is just a matter of the products already existing and not getting traction (and I am not talking about the other big search engines).

I just learned of your site through this post, it looks useful (and not just for SEO spam), thank you.

Some feedback: some initial searches did not find blogs that were of interest. When browsing I noticed the tags. I wish there was a page of tags that I could browse. If there is, I could not find it.

Also many of the site titles do not provide useful information on their topics and are untagged. If you could auto generate some tags, perhaps based on word frequency in post titles (with some stopwords no doubt), that would improve the browsing experience, at least for me.

I am slowly convincing my coworkers that deploying the exact same binary as two different 'services' is a significant tool to have in your toolbox. Some disaster recovery work we're doing is making it a much easier sell.

I'm really just combining two very old tricks here. Traffic shaping based on class of service for two different requests, and for two different classes of users.

Segregating bot traffic improves consumer experience. Segregating admin traffic from both allows you to set an upper and lower bound on availability.