back

by m-i-l·3y ago·view on hn ↗
There are a couple of challenges with the word vectorization + cosine distance approach:

1. The current embedding models are for relatively small amounts of content rather than whole web pages. For example, most of the popular Huggingface sentence-transformers are for 256 tokens max, which is only sufficient for a small part of most web pages. There are some models for larger sizes, e.g. Longformer and BigBird, but they're still only 4096 tokens, and there aren't any ready-made implementations for cosine or dot product similarity at the moment so you'd need to roll your own. Some of the new LLM OpenAI embedding models do go to 8191 tokens and can be used for cosine similarity, and costs may be acceptable for the content embedding, but for the query embedding you'd have to be super-careful about costs because if you put any kind of search on the public internet you could find that you'll soon be overwhelmed by SEO spam bots (most of which even Cloudflare can't block) and that could become prohibitively expensive to service via a paid-for API solution like OpenAI. A popular approach to the token length issue is to chunk your pages into multiple blocks.

2. Even if you get the embedding and chunking approach working, that is still just for page comparison, and single web pages are not necessarily representative of whole web sites. For example, if comparing home pages, many blogs just have a link to the posts page on the home page without any of the actual content from the posts. Or some websites cover an eclectic range of topics, and if your goal is to find other sites that cover a similar eclectic range of topics then you you need an embedding for the whole site rather than individual pages. Not sure about the best way to address this one though. Apparently averaging embeddings may work surprisingly well but I haven't tried that out yet. Other possible options might include summarising (chunked) pages and summarising the summaries for the site, or using topic modelling (e.g. BERTopic) alongside some kind of website taxonomy, or something like that. Keen to read of other possible approaches here.