Yeh it’s weird they don’t mention parameter size or other reasoning metrics. It’s a very cool approach to getting structured output from an LLM, but the benchmarks don’t show us the whole picture. I’m wondering if their approach can be used to delegate to different models at each step in a structured output. If it could be run with mistral 8x7B and still maintain its performance then that’s awesome.
back
1 comments
The last bit is already under way with speculative decoding but wonder what exactly are they proposing