back

by numeri·2y ago·view on hn ↗
The linked leaderboard is actually very trustworthy, in that it consists not of scores on a test dataset, but of ELO ratings generated by actual humans' ratings of the models' responses.

You can go and enter any prompt you like, wait a bit, and then get two LLM responses back, which you can then rank or mark as tied, after which you'll be shown which model each came from. Maybe you already knew this, maybe you didn't. In any case, I don't see any real way for Goodhart's Law to apply here – the metric and the goal are the same here, i.e., human approval of answers.

2 comments
The linked leaderboard is not at all trustworthy to generally rank LLMs (i.e. what everyone uses it for) because the sample bias is absurd. It ranks LLMs by how they respond to queries that users of the leaderboard are likely to test on. Which is about as representative of general usage as the average user of the website is representative of global society (i.e. in no way whatsoever).
The combined rankings on HF aren't just the arena scores, and certainly MMLU is an example of Goodhart's Law at this point.

The Chatbot arena is more an issue of sampling bias, and I think it would be pretty interesting to run an analysis of random samples on the prompts provided to see just how broad they are or aren't.

It is trivially easy to see just how massive the sampling bias is. There are countless tasks where e.g. the gap between a version of Mistral and a version of GPT are incredibly large. Anything that requires less common knowledge, anything creative in a random language, let's say Bulgarian, or Thai. Yet on the leaderboard things are very different, because the arena is only used by developers who verify performance on tech-related prompts in English.