back

by artninja1988·1y ago·view on hn ↗
>Ultimately, R1-Zero demonstrates the prototype of a potential scaling regime with zero human bottlenecks – even in the training data acquisition itself.

I would like this to be true, but doesn't the way they're doing RL also require tons of human data?

1 comments
I think yes. But hopefully in math with compute advances we can lower the human data input by increasing the gap that is bridged by raw model capabilities vs search augmentation (either with tree search or full rollouts)