I've created a few similar tools for link scraping: https://github.com/chapmanjacobd/library#usage
- library links-extract: extract inner links from pages (stdin, local files, or remote sites)
- library links-add: extract inner links across from multiple pages (paged lists of articles/forums) into a SQLITE database
- library links-update: fetch new items for each link in a links-db (just added this yesterday)
You can use the same filtering across all subcommands. For example you can filter based on text between links:
pip install xklb
library links https://en.wikipedia.org/wiki/List_of_bacon_dishes --path-include https://en.wikipedia.org/wiki/ --after-include famous
You can use all of them with the same --cookies-from-browser interface that yt-dlp provides for pages that are locked behind a login.Alternatively, copy and paste:
library links --local-html <(xclip -selection clipboard -t text/html) --after-exclude paranormal spooky horror podcast