back

by tosh·11d ago·view on hn ↗
it's very difficult and time-consuming to make good comparisons (especially for open ended tasks) also because there are so many configuration options

  - which env do you provide?
  - which model(s)?
  - subagents?
  - system prompt (default or custom?)?
  - agents.md
etc etc

also you kinda have to look at many runs and study their traces, if you look at too few runs the outcome variability you're drawing from is too high