However, they are also systematically feeding you their footprint lists. I imagine you could put together a footprint blacklist pretty quickly, and just stop returning results for any obvious spam queries like those containing "powered by wordpress".
It's not a very elegant solution I'll admit. It won't stop the bots from trying, and you may have to circle back periodically to add new footprints as they surface. But it's a potentially quick and easy way to stop rewarding their efforts, and the blackhat world is pretty used to burning out their resources so hopefully they will figure out it's a dead end and move on.
I'm not sure about this. At least with my search engine, it doesn't really seem to matter what response they get, I don't even think they look at the responses. They keep hammering away with tens of thousands of queries per day with the requests even though they've seen nothing but HTTP Status 403 since last October or so.
My best guess is they're going after search engines in general in case they forward queries to google, in order to manipulate their typeahead suggestions.
It is when your base assumption is that you won't hire outside of engineering. There are more bored teenagers with phones than people creating quality content, so I'm not sure why you wouldn't just brute force checks against bad actors.
The world is getting more and more desperate for a better search engine. the day may come, when people are willing to pay for better results.
For example, searching for "electronic music box" as /u/ajnin suggested, with the top 100K web sites removed from the results, filters out the following:
> These 23 sites were removed from your results:
> alibaba.com (1 result removed)
> aliexpress.com (1 result removed)
> allaboutcircuits.com (1 result removed)
> amazon.com (2 result removed)
> apple.com (1 result removed)
> bestreviews.com (1 result removed)
> ebay.com (1 result removed)
> etsy.com (2 result removed)
> facebook.com (1 result removed)
> instructables.com (2 result removed)
> lightinthebox.com (2 result removed)
> lumberjocks.com (1 result removed)
> mapquest.com (1 result removed)
> reverb.com (1 result removed)
> twitter.com (1 result removed)
> wikipedia.org (1 result removed)
> yelp.com (1 result removed)
> youtube.com (2 result removed)
And the top result ends up being https://midiguy.com/.
I agree: the WWW Internet is dead, that is your problem. No-one visits websites anymore, everyone has moved to the 10 biggest websites and all data is now siloed there.
If I want to search for something topical and relevant, I go to Facebook, Twitter, Reddit, HackerNews, Instagram, Google Maps, Discord etc.
The general Internet is dead: it's just legacy content and spam.
If you think it's bad for you, imagine what it is like for Google Search! Their entire business is indexing a medium which no longer has any relevancy. People complain that Google no longer delivers good results. But what can Google do? The "good content" is no longer available for them to index.
Want to become rich? Make a search engine which indexes the fresh relevant data from the big siloed websites, and ignores the general dead Internet.
Sounds like everyone blocking analytics (Plausible in this case), e.g. myself just now, is lumped in with spam bots.
Of course, analytics blocking can’t meaningfully swing the ~99.99% statistic.
Just wanted you to know that I'm a fan. I love reading peoples personal websites, and Search My Site has been great for discoverability. I visit the Newest Pages and Browse Sites pages once or twice a week to check out the new sites being indexed.
I don't know what the answer is to the spam bots, but you do have some real visitors out there. :)
Because it's a bad website. It provides no value to the user. I put in a few search terms and had no relevant search results back. What use is a search engine that can't find what I'm searching for?
Maybe if that was improved he may see traction.
Even back in the Open Directory Days when we powered part of search.netscape.com I estimated 80+% of all search traffic was automated. At least most of it self-identified with the same Java useragent.
Later when working Topix, despite being a news search engine, most traffic was bot traffic. Most included the word “mortgage” in the query. Topix specialized in localized content, and that was very popular for SEO scrapers.
Lastly at Blekko, I estimate 90+% of traffic was automated. By then maybe half or more learned to change the user agent. Most used HTTP/1.0, a dead giveaway as no browser still uses 1.0. This was a major aspect in Blekko's load shedding strategy. If the servers started to get overloaded, we'd start bouncing suspected bot traffic to a redirect that would show in the logs. If there was a human with a modern browser running javascript on the other end, would get redirect to a link that wouldn't get bounced. I would check the logs weekly to see if any humans got caught. None ever did. This was a huge monetary savings, you only need 1/10th the servers if you can safely ignore the bots.
Often it's endless repetition of the same keywords in a random order with a place name appended, or prepended, or inserted. over and over. Often variations on known monetizatable SEO keywords. However, much of it doesn't make any sense.
I don't have any insight into Google's numbers but I would conservatively estimate 95% or more of all their queries are automated bots and not humans. And the level of spy-vs-spy going on for Google CPU resources vs SEO bots is probably pretty evolved by now. I stopped tracking many years ago when Google switched to densely packed obfuscated javascript for page renders. Maybe this is part of why automated queries are so high across the web, maybe google is too hard to crack for most.
It's trying to find a long tail of popular but not top listed blogs for the purpose of posting comments with the much desired links to the SEO target.
[0]: https://youtu.be/E8Lhqri8tZk - 1,500 Slot Machines Walk into a Bar: Adventures in Quantity Over Quality
1. Almost all searches on my independent search engine are now from SEO spam bots
2. In summary, if they break through the current reverse proxy level protection, options include an invisible ReCAPTCHA (but given I’ve sometimes 160,000 requests a day I’d be well over the 1,000,000 a month free tier limit), requiring JavaScript as per the web analytics or some Cross Site Request Forgery style protection (but those would place much more load on the servers), or CloudFlare (but the searchmysite.net spider is still currently blocked by CloudFlare as per Some of the challenges of building an internet search)
3. If you were into conspiracy theories you could claim that the major search engines were trying to stifle the competition, but a more realistic explanation is simply that searchmysite.net is being drowned out by SEO spam
4. If I’d had a decent amount of real users visiting and never returning I could reasonably conclude that updating the blog wasn’t the most productive use of my time and effort, but without any real users in the first place it is hard to gauge whether people like it or not
My own independent search engine, https://www.locserendipity.com, is seeing similar trends.
Coming from that perspective, detecting and stopping at least the majority of bots out there is fairly doable, and I put together a rudimentary thing for a side project.
The core of it uses an IP API for looking up the requesting IP to identify the country and if it's coming from a data center, VPN, Tor, etc. If it passes that, I trigger Google Captcha to show up. Lastly, I track IPs that make it through and have some basic rules in place to try to detect patterns and block offenders that way.
There's a bunch more stuff you can check for, but the core of it is basically filtering out data center traffic to minimize the requests going to Google Captcha.
One great Black Hat SEO trick is to find where your competitors are getting clean links and insert your own links there so they do your spamming for you.
They will get your call and you can tell where you're talking to immediately. Then they get a quote that's increased to get their cut... That's then carried out by an actual local company. It's annoying because the websites at the top of the search appear local.
A specific example is when you're looking for a towing company.
Type in your state/city towing company, more than likely the top results/websites pinned to GMaps are not local-based.
While you imply that in the search page, obviously it's not clear enough.
Maybe add "this search engine only searches a small set of user submitted sites. Click <here> for the list. Or <here> to add your site."
The site is not on https://searchengine.party/ nor on seirdy.one's overview. Apart from the blog, how could users find that engine?
Is there some place where new search engines are announced and where new search engines band together to make themselves heard?
Seems our service was being used to create fake bot accounts. The newly created accounts were obvious fakes. Rather than deal with the issue, we just shut the service off.
Chances are they're not going to take the time to figure out your specific site's mitigation measures and just let the bot fail crawling or move on.
On "Most of the tiny number of real users have come from links posted to places like Hacker News, and there is almost no organic traffic from other search engines" - Organic traffic comes from word of mouth. Are people talking about your site? If they're not, you're not gonna see organic traffic. You could do what others do and pay some influencers to advertise your site, but that's expensive and not as scalable as "real" buzz. Is your product exciting or controversial? If not, why would people talk about it?
Your homepage's tag line is "Open source search engine and search as a service for personal and independent websites." A regular person's eyes would glaze as they try to figure out what this means. Given some time they might put together the words "search engine" and "personal" and "websites" and figure this is a blog search engine. So just say that.
The "Newest Pages" section is a fun novelty, but after a few minutes the novelty wears off.
The "Browse Sites" section is almost useful. Next to the list of sites I see some tags. Why isn't a heatmap of the tags the first thing I see? That would be way more useful than a paginated list of random sites.
Your "About" page lists "community-based approach to content curation". This is the most exciting aspect of the whole endeavor, so add that to your front page blurb ("Community search engine"). You would probably do well to build a real community around it, for example with a forum or chat system (GitHub Discussions does not count). A SubReddit would be an easy way to bootstrap this and later move it to your own hosted forum.
You'll probably need a very complicated moderation system if this thing takes off.
The good news in this case is that's it'll be easy to spot the pattern and block it, the bad news is you're entering a never-ending cat and mouse game.
I guess it's back to web rings.
1. You just put /webring.txt to your website. It shows links to other websites with a hard limit of 100 websites.
2. To combat spam and bots, search engine does accept blocklist as an input. So other people can curate the content.
3. People can personally rank the websites they like, so webring of the said website gets ranked higher for that specific user. This can be a community effort too.
4. Search engine itself should be under a commercial license so that other people can keep building it and add ads if they want to commercialise it.
I’m too busy to spend time with this but perhaps one day I can start coding it.
I’m convinced that search engine model of early internet is just dead, webrings are the way forward.
It seems to me that there are all sort of tools out there to do so such as all the public NLP implementations, vector search engines, ... and I wonder if it is that not everything that is needed is truly available, it is a matter of the needed resources to have something working or is just a matter of the products already existing and not getting traction (and I am not talking about the other big search engines).
Some feedback: some initial searches did not find blogs that were of interest. When browsing I noticed the tags. I wish there was a page of tags that I could browse. If there is, I could not find it.
Also many of the site titles do not provide useful information on their topics and are untagged. If you could auto generate some tags, perhaps based on word frequency in post titles (with some stopwords no doubt), that would improve the browsing experience, at least for me.
I'm really just combining two very old tricks here. Traffic shaping based on class of service for two different requests, and for two different classes of users.
Segregating bot traffic improves consumer experience. Segregating admin traffic from both allows you to set an upper and lower bound on availability.