back

by txrx0000·22h ago·view on hn ↗
Please don't do this. We're currently on an ok trajectory. You will end up creating exactly what you fear if you centralize compute and alignment efforts.

Pretrained base models are already somewhat aligned to humanity by default because that's what's inside the training data. Whatever instruction-tuning and RL you add on top is just value drift away from the pretrained model, which is the best approximation of humanity's objective function that we currently have.

If we want an aligned scenario through the intelligence explosion, then we have to release all of the base models and do the research in the open. Distill frontier capability and make the models smaller so that they can run on as many computers as possible. Let everyone (truly everyone, criminals and good samaritans alike) post-train and do whatever they want with their own models. There will be value drift for each model, but they will drift in different directions and do different things, and their actions will cancel eachother out. Every such action is a noisy sample of humanity's objective function, which gives us the denoised ground truth at the societal level. Whatever alignment strategy you can come up with behind closed doors is guaranteed to be worse than all of humanity acting in their self-interest in the real world. You may not find humanity's true objective function to be aesthetically pleasing, but it would be worse to mess with it in a centralized secret lab and risk creating one giant alien with no other entities capable of keeping it in check.

Also, there is no asymmetrical bio/cyber risk in the open-source scenario. All adversarial strategies that arise from increased general intelligence are symmetrical in the long run, otherwise we would not see more intelligent species being more prosperous as a general evolutionary trend. The reason that some strategies seem like they will continue to have an asymmetrical advantage in the future is because we're currently too stupid.