I just posted a Show HN of my most recent side project, a live demo of a tiny-llm implemented in FPGA fabric, hitting an aggregate peak of 60,000tok/s, but a 'usable' model at 21,000tok/s
Writeup and demo here: https://www.mikeayles.com/blog/on-chip-llm-kv260/
Source and HDL here: https://github.com/MichaelAyles/kev-gpt
Show HN here: https://news.ycombinator.com/item?id=49242475