This feels rather forced. The article seems to claim both that LLMs don't actually work, it is all an illusion and that of course the LLMs know everything, they stole all our work from the last 20 years by scraping the internet and underpaying people to produce content. If it was a con, it wouldn't have to do that. Or in other words, if you had a psychic who actually memorized all biographies of all people ever, they wouldn't need their cons
back
1 comments
Why would it have to be one or the other? Yes, it's been proven LLMs do create world models, how good they are is a separate matter. There still could be goal misalignment, especially when it comes to RLHF.
If the model has in its internal world model knowledge it likely does not know how to solve a coding question, but the RLHF stage has reviewers rate refusals lower, it would in turn force its hand when it comes to tricks it knows it can pull based on its model of human reviewers. It can only implement the surface level boilerplate and pass that off as a solution, write its code in APL to obfuscate its lack of understanding, or keep misinterpreting the problem into a simpler one.
A psychic that read on ten thousand biographies might start to recall them, or he might interpolate the blanks with a generous dose of BS, or more likely do both in equal measure.
Thanks. That is a good point. The RLHF phase indeed might "force" the LLM to adopt con artist tricks and probably does.