We've been approaching it a little bit differently. We think larger more capable models would actually immediately improve the performance of Skyvern. For example, if you run it with LLaVa, the performance significantly degrades, likely because of the coupling
But since we use GPT-4V, and it's rumoured to be a MoE model, I wonder if there's implicit decoupling going on.
I'm gonna spend some more time thinking about this