back

by Alifatisk·1mo ago·view on hn ↗
> There were some tests last year-ish from hf that showed that simply alternating (randomly) between claude and gpt (whatever their versions were at the time) on a task produced better results than either of them individually. So during a task, the first call was sent to one, then the other and so on.

Where can I read more about these tests?