back

by wwilson·2y ago·view on hn ↗
A drawback of our approach is absolutely that it is expensive to test extremely large volumes of data or compute this way. Even before you start running into physical limitations of our current platform, you will probably be complaining about your bills. :-)

Our advice on this is that there are actually a lot of things you can do to exercise behaviors that are usually only seen at massive scale. For example, if you run a distributed storage system, you can probably configure it to split and move shards at 1/1,000,000th of the production size. That might let us hit a tricky codepath much more cheaply. We have a lot more about this in our documentation, e.g. here: https://antithesis.com/docs/best_practices/optimizing.html#k... and here: https://antithesis.com/docs/best_practices/find_more_bugs.ht...

The other thing is just that the reason many bugs only happen at scale is that they're some kind of subtle distributed race, and you need a lot of nodes for one of the runners in the race to be slow enough that the other sometimes wins. But we can very easily and efficiently provoke these sorts of races by pausing individual threads or freezing nodes, etc. We actually pretty regularly hit issues with tiny deployments that our customers only see in their largest clusters (but no promises, this obviously depends on the details of the software).