it's very difficult and time-consuming to make good comparisons (especially for open ended tasks) also because there are so many configuration options
- which env do you provide?
- which model(s)?
- subagents?
- system prompt (default or custom?)?
- agents.md
etc etcalso you kinda have to look at many runs and study their traces, if you look at too few runs the outcome variability you're drawing from is too high