Thank you. This has been in my mind for past 1 year. Wanted to do it using vector embedding similarity match, but due to costs and compute requirements, had to resort to keyword based.
The data comes from daily-updated public BigQuery dataset: https://news.ycombinator.com/item?id=40644563