>Ultimately, R1-Zero demonstrates the prototype of a potential scaling regime with zero human bottlenecks – even in the training data acquisition itself.
I would like this to be true, but doesn't the way they're doing RL also require tons of human data?