back

by moultano·5y ago·view on hn ↗
I'd bet that a shallow transformer that was distilled on gpt2, possibly with quantized weights, would easily win. GPT2 isn't optimized for size at all, but there are tons of ways to do so.