back
79 comments
My personal preference is to 'ip route add blackhole ${net}' as it has the lowest CPU overhead and I can add hundreds of thousands of CIDR blocks with no noticeable impact. The only downside is that it won't stop UDP packets from getting to a UDP listener. There will not be a response but the application will still see it. For my TCP daemons it's great.

    grep -m1 -E ^Tot /proc/net/fib_triestat ;ip route | grep -Fc blackhole
    Total size: 56735  kB
    426951
Those 426951 blackhole routes include data-centers, VPS providers, botnets, AI datacenters that ignore robots.txt, search engines, abused CDN's, known bad residential nodes and much more. I still see a few residential proxy bots that do a halfway decent job of pretending to be real people at times but the feds are playing whack-a-mole with them. The bots self report to my silly blog so I can block them elsewhere on systems I might care a little bit about. Happy to share them if anyone is remotely interested.

I also use a couple generalized rules in nftables raw table that keeps a lot of beyond poorly written bots away including hping3 tcp floods and masscan. My rules to port 443 are stateless. One must not taunt the state table.

The problem is when you block those "residential proxy bots" you also block real people who just happen to have a dumb teenager on their network playing some free games that are monetized by proxies.

The only real solution to bots is making users log in. And even then you have to fight registration bots.

I get what you mean. It happens all the time when some clown trashes an IP's reputation and Cloudflare or Google will send the next lease of that IP into crosswalk fire-hydrant bus traffic light purgatory.

That's why I eventually let those go usually after a kernel update and the git repo for FireHOL gets updated often. The kernels get updated often. I only perma-ban the data-centers which is fine for my silly blog and probably for some peoples hobby sites. People can chose which methods to apply, how to apply them or which ones to skip entirely.

Excellent username btw. Those SNL Celebrity Jeopardy episodes are unforgettable. [1]

[1] - https://www.youtube.com/watch?v=bEghu90QJH4

I would like to learn more how you maintain your table of IP ranges (or CIDR block). How do you decide when to add/remove a range?

I'm most concerned about blocking innocent users, currently I use Cloudflare to block known bad ASNs using a list I found on GitHub.

How do you decide when to add/remove a range?

The only IP's that come and go are the Tor 30 day blocklist and a couple FireHOL attackers from a repo though I will sometimes leave the last entries live until reboot. I do not really need to block tor but I use this silly blog as a testing ground. Tor and some known abusers come from a git repo I refresh periodically.

The data-centers, VPS providers, CDNs, known botnets are perma-banned. For my hobby nodes I personally find this acceptable. I would not do this in a professionally managed data-center. There are better methods for those cases especially for B2B corporate arrangements. Regardless of what daemons I run I never have external dependencies that need to be accessed from my node or from the client with exception of stratum-1 time servers.

I do have to periodically update the CIDR blocks for given ASN's. I have not automated this but I probably should some day. It's not hard to automate, I am just excessively "efficient". I was told to stop calling myself lazy, but I am.

Methods 2, 3 and 5 are the ones I talk about here. [1]

[1] - https://nochan.net/b/Internet-Crap/20260606-How-To-Block-Som...

What does all of this give you? For a static(?) site burning a few cycles unnecessarily, saving what, 30 cents of power per year?

Peace of mind? Fair enough but I'd be more wary about blocking legitimate users. VPS providers are often used for VPNs etc.

What does all of this give you?

Good question. A playground to test things. A place for bots and their kin to report themselves to me so that I can use this information for sites I actually want to protect a little bit. A place to share some ideas with a small handful of like minded people. I do not consider power savings for a device unless I have it running on one of my inverters or if I am doing that to constrain a potentially malicious node.

Blogs are throw away for me. After some time I delete the VM and edit articles offline for 6 to 18 months and then put them back up on another domain when there may be a need to share some old articles. This method disjoints the archive sites and breaks any filters botters have set up to ignore me. That also allows me to change the CSS. I try to make it smaller each time.

I also find it easier to put long form content on a blog of sorts instead of HN comments in the unlikely chance that YC removes HN due to future ID/Age verification constrains or other unforeseen reasons that we hope never happens.

> The only downside is that it won't stop UDP packets from getting to a UDP listener. There will not be a response but the application will still see it.

Try:

  ip route add src ${net} blackhole
You could configure reaction to add and remove `ip route` commands in that format.
> This software is gay, trans and anticolonialist. If you're uncomfortable with that, please don't use it

Weird message to include in AGPLv3 licensed software (which explicitly allows people to use software however they like, regardless of their beliefs or feelings).

You can have preferences while not restricting legal rights.
You can, but if the exact quote in the GP is correct the claim is claiming the software is "gay, trans and anti-colonialist" and asks you not to use it. Why use a license that is designed to be politically neutral and then ask some people not to use it?

What I can see is a fairly clear indication that they do not want contributions from people whose politics differ from theirs. I would also question whether government funding of a project with political policies about who can participate is appropriate. The political stance is also rooted in a particular culture so is unwelcoming to people from other cultures.

Of course people can political views and preferences, but they presumably have some aim in mind when making that statement in the README. What is that aim?

A license designed to be politically neutral?

The GPL variants are the antithesis of politically neutral.

>Why use a license that is designed to be politically neutral and then ask some people not to use it?

Because you can have preferences while not restricting legal rights.

> What I can see is a fairly clear indication that they do not want contributions from people whose politics differ from theirs

This is the same FSF that in the past has refused contributions from people whose politics include "I would like this software to run on my windows/apple/other proprietary platform". They're extremely political.

1. GPL is extremely and intentionally political 2. anti-colonialism is not a trait of culture 3. the state of being gay/trans is not political 4. culture != politics 5. is != ought

Of course people can have their preferences, but presumably that reaction implies "gay/trans anti-colonialist" ought to be a problem for some "political/cultural" groups?

What is the aim of that?

My read is that the authors are signalling their community is a "safe space". I don't think they're trying to exclude particular cultures, trigger people, or otherwise cause problems.

You can always ask anyone not to use anything. Doesn't mean they have to listen, but you can still ask.
Because signalling has replaced real virtue.
My Gen Y brain can't read the phrase "[inanimate object] is gay" without interpreting it as disapproval.
My Gen Y brain can do this no problem, and I'm among the oldest of the generation. Anthropomorphizing is fun!
Functions are colored. Software has sexuality. And we aren't even talking about AI!
That's interesting. I haven't used fail2ban for a long time, but reaction is worth evaluating. Unfortunately, that post does not describe their full configuration. Maybe it's on purpose, so that attackers can't adjust to fit.

My experience is that modern web scraping had no obvious pattern, since it is proxied through many IPs. The last time a server was failing to handle the pressure, we decided to temporarily ban IPs from some Asian regions. How does the FSF decide to ban an IP?

Why do they use iptables + ipset instead of nftables? Is there a technical reason or is it just legacy? AFAIK, Nftables is more performant, and IMO simpler. And it has native sets, see https://wiki.nftables.org/wiki-nftables/index.php/Sets

> The last time a server was failing to handle the pressure, we decided to temporarily ban IPs from some Asian regions.

This is something we've been forced to do at work, a LOT. Some weeks it's Huawei Cloud, Tencent, and Alibaba. Other weeks it's all China Telecom. We're using Anubis where possible, but a lot of it is just whack-a-mole with residential proxies. I looked at Datadome and HUMAN, but they would be hundreds of thousands a year at our traffic scale, and I suspect may also have false positives. We abandoned CrowdSec for that reason as well.

I'd love to find a decent k8s native solution to this problem.

Prosopo could cut out a lot of the residential proxy nonsense for you. We integrate with lambda@Edge / cloudfront workers / server side and perform analytics to detect residential proxy networks - at far less cost than DD or HUMAN.
> I looked at Datadome and HUMAN, but they would be hundreds of thousands a year at our traffic scale, and I suspect may also have false positives.

I can confirm. DataDome has been making my life a living hell. It thinks my phone is a bot, so I can no longer use PayPal and other quasi-monopoly services while on the go. DataDome will reliably block me on the first request.

Fuck DataDome.

iptables has been mostly a wrapper for nftables for some time now. The choice of iptables + ipset with reaction is the difference in their configuration. Compare restart performance between the ipset and nftables example configurations with lists of greater than 1 million IPs.
It's somewhat interesting to see the FSF's approach to this. From what I understand they can't really use something like anubis since they want their websites to be accessible without javascript:

https://www.gnu.org/philosophy/javascript-trap.html

Users can't consent to running a page's javascript the way they can consent to running a program they've intentionally downloaded, so it's effectively "non-free" regardless of license.

Anubis does support the no-JS HTTP meta-redirect proof of work but few know about it and fewer enable it. And it may not block everything.
I indeed did not know about this. There seem to be some caveats:

https://anubis.techaro.lol/docs/admin/configuration/challeng...

I guess for the meta refresh challenge it's less "proof of work" and more "proof of patience".

Does it still need a cookie though? Another thing I have disabled by default.
For what it's worth, Anubis supports LibreJS: https://github.com/TecharoHQ/anubis/blob/main/web/build.sh#L...
> We placed our regular expressions in fail2ban, and found that we were hitting the maximum rules that could be added to UFW firewall rules on our systems which showed degradation around 65,000 rules

Firewalld had a similar issue up until recently as well.

>Popa botnet

It's no more of a botnet than ProtonVPN for example. Apps intentionally added the Popa SDK to their apps as a monetization method. This allows apps without ads and tracking to be financially viable. I would expect FSF to support apps being able to move off of monetization schemes that depend on tracking people so it is disappointing for them to put such alternative monetization technologies in a negative light.

This monetization scheme benefits the botnet controller and the developer who added the SDK and not the user who likely did not realize they signed up to become an exit node.
It allows as free versions of apps to be economically viable and compete with others. It helps users because they don't need to be spied on and shown ads to fund the development of the app.

The existence of an app brings users value, else they wouldn't use it.

"Monetization". What a horror.

I pay for some software services. The services I pay for have a billing page (or a donation page) and I pay via the banking system

I rigorously block every ad, every tracker, every thing that does "monetization"

The evil period of trying sneaky ways to generate money is, I am optimistic, coming to an end.

If you want my money, ask me. If you must have my money, demand it. If you are sneaking around "monetization" I will do everything I can to stop you.

Yeah but they don’t want your money. They want the botnet users money.

You got a genuinely free, to you, app.

"Many sysadmins know about fail2ban..." and many will now know about reaction. But why will the result be any different than fail2ban? It won't.

I identify features (which can be expressed as firewall rules) from log data; I write totals to a temporary store (Redis). I have periodic tasks which scan the temp store for patterns which exceed thresholds. When that occurs, fail2ban creates the appropriate rules. This occurs in depth and in concentric rings.

Et tu?

The difference between fail2ban and reaction is performance. If you are not hitting the ceiling of fail2ban, then you may not need reaction.

Do you have a blog post about your automated fail2ban rule generation?

Why should I have a blog post? I get amazing "performance" out of fail2ban, it does what it's supposed to do and I don't ask it to do more. I've given it "super powers".

Last year I blocked basically all of Brazil. No problems. Who knew?

are scrapers attackers?

I get they're DDoS; but take the mask off, and arn't they just the AI monied interests that fund the FSF? and a lot of them are just active inference, eg, the user is trying to ask about something and the AI monied interests setup a web scraper to go and get that data.

Just seems like no one wants to call out the hand that feeds them in a human centipede that's best described as the torment nexus.

instead of scraping then, they could pay the fsf for a dump of the site or some API access or something, right? why overload the servers normally.
> AI monied interests that fund the FSF

Can you elaborate on who these interests are precisely?

Anything which fills my logs with garbage is unwanted. Your cat could have fallen asleep on the keyboard, I don't care. If you want to use the internet as a giant petri dish, that's on you; but the cat box is elsewhere. I can feed you garbage or block you because your code is shit, or I don't like your style.