Actually Nutch is used to produce the Common Crawl[0] and 60% of GPT-3's training data was Common Crawl[1], so in a way it is being used by a lot of web users today.
I did very briefly look into the possibility of self-hosting a search based on the Common Crawl data, but the latest size is 390Tb, and given the largest consumer drives are around 10-14Tb, that would be a lot of drives and a lot of money. Plus by all accounts most of Common Crawl is poor quality spam type of content, which OpenAI, Google etc. try to filter out before using it to train their LLMs. That's why I don't think the solution is a search engine for the whole internet any more, but one which just indexes the relatively small minority of good bits.