back

by numeri·2y ago·view on hn ↗
You're not necessarily wrong, but I'd imagine this is almost prohibitively slow. Also, this model seems to use two experts per token.