That's not to belittle the considerable achievements of Ladybird; their progress is really impressive, and if web-platform-tests are helping their engineering efforts I consider that a win. New implementations of the web platform, including Ladybird, Servo, and Flow, are exciting to see.
However, web-platform-tests specifically decided to optimise for being a useful engineering tool rather than being a good metric. That means there's no real attempt to balance the testsuite across the platform; for example a surprising fraction of the overall test count is encoding tests because they're easy to generate, not because it's an especially hard problem in browser development.
We've also consciously wanted to ensure that contributing tests is low friction, both technically and socially, in order that people don't feel inclined to withhold useful tests. Again that's not the tradeoff you make for a good metric, but is the right one for a good engineering resource.
The Interop Project is designed with different tradeoffs in mind, and overcomes some of these problems by selecting a subsets of tests which are broadly agreed to represent a useful level of coverage of an important feature. But unfortunately the current setup is designed for engines that are already implementing enough feature to be usable as general purpose web-browsers.
PS I'm a big fan of the work and appreciate what you do. I check the interop page about once a week!
Still an amazing feat of development from the entire team.
There’s still a very long way before they can compete with Chrome, of course. And I’m not sure I ever understood the value proposition compared to forking an existing engine.
Though, I suppose even if true, it would still be a pretty good timeframe.
Just building a good html/css renderer and a JS engine is crazy, but now you are hooked into the ecosystem and at the mercy of whatever comes next. Chrome can push back against proposals but little browsers either use chromium or are basically in a riptide trying to make sure they keep up.
And in that sense, is it better than Gecko with firefox, which is non-profit?
Well, it could be that AI actually speeds up development, who knows.
The real test isn't passing 90%—it's whether they can keep pace as the web platform adds new APIs faster than any independent team can implement them. Browser engine development has become a regulatory moat, and breaking it requires either massive funding or accepting permanent incompatibility with the "modern web."
Still rooting for them. Browser monoculture is worse than metric gaming.
I suppose their success is likely directly related to the fact they made reasonable, practical development choices, but still.
Me as customer: oh man I'm sure glad stuff is reviewed to some quality bar and the OS limits API access.
"Oh, is this metric important? Let me get right on that."
No shade intended towards the Ladybird team. You were given the terms and you're behaving rationally in response to them. More power to you. It's just a fantastic demonstration of what it looks like to very suddenly be developing against a very specific metric.