back
100 comments
Article from April OP;

Some discussion then: https://news.ycombinator.com/item?id=47633506

Measuring success in LoC is so wrong… It's like bragging that you took 10,000 steps to reach the store by walking in circles when you could have taken 500 steps if you just walked straight. The end result is the same.

AI has a tendency to generate more code than necessary. It keeps re-inventing things, and every time you ask it to add a new feature or fix something, it just keeps on piling the code. I now periodically ask AI to refactor the code by simplifying, removing unused things, factoring out, and reusing.

> It's like bragging that you took 10,000 steps to reach the store by walking in circles when you could have taken 500 steps if you just walked straight.

My wife does exactly that. For the exercise. She makes large detours to go anywhere. The end result is a healthier body.

How this translates to software, I don't know. I don't think AI benefits from this exercise.

> It's like bragging that you took 10,000 steps to reach the store by walking in circles when you could have taken 500 steps if you just walked straight. The end result is the same.

The case is different with LoC though: the more the worse. I'd much maintain the thing written in 500 lines than the one written in 10,000 lines.

IOW, you're doing an insane amount of micromanagement with a most inexperienced beginner developer.
"found numerous examples of bloat and inefficiencies in Tan’s site code, and used a single (Anthropic) Claude session to review the files he downloaded from the website to confirm his observations"

1. I hope they never get hold of the code of MS Office or almost any other piece of real-world business software.

2. So anyone with claude access could arrive at the same conclusions ... and ask claude to fix it?

Personally, I've used my newfound productivity to build higher quality software at the same rate as I produced mediocre quality software before. Some people, like Gary Tan for example, seem to instead produce software at the same quality or worse than before, but just produce a ton more software regardless of quality.

I guess the article in reality is just these two perspectives pitted against each other for some cheap views.

Yes but for that you need to focus on the quality. There are probably horrible issues on backend as well.
I'd suggest looking at the review itself, there's an X-the-everything-app thread on it.

https://x.com/Gregorein/status/2038953944475472316

Note that Rails was built as a framework for making blogs, I'm having trouble understanding what 78,000 lines of ruby in the context of a Rails blog could ... do.

I'm sure there's some powerful ugly stuff in Office but in a good code that's calcified kind of way. It got that way over like 30 years of releasing to the public across platforms, not over a weekend.

I'd be surprised if microsoft.com is shipping their entire test suite unminified and their back-end posting rich text editor with index.html (with two title tags in the head) and rendering the entire DOM for desktop and mobile regardless of your platform.

I'm not critiquing Garry or the site. I think it's great people are using AI to build things that bring them joy, or that they find useful. I certainly do.

I am opposed to the idea that we've decided to go back to measuring work in terms of lines of code. It has always been the worst metric on earth as a proxy for productivity. Every line is a liability, and it always was. AI has not changed that, if anything it's amplifying it.

The best PRs remove code, not add, and the only companies that seem to have exponentially grown their revenues in line with AI-generated LOC are OpenAI and Anthropic. Everyone else seems to be rummaging around for an ROI.

> I hope they never get hold of the code of MS Office or almost any other piece of real-world business software.

Except he's not building Office, a software with decades of legacy, used by hundreds of millions of users. He's coding a website, effectively writing bloat with a silver lining of useful features. AI automated and inflated the worst of practices too. Anyone outputting 37K LoC daily is creating bloat and inefficiencies at unprecedented rate.

And enterprise software was already the butt of all jokes in this regard. We found ways to make that worse at scale even for the basic things. It's not a good look when you need to use this "whatabout".

That was the strangest line for me as well. The article tries to play up this "Polish engineer with an MSc in computer science who has 13 years of industry experience" who ... used claude to review the code. My mom could have done this without an MSc or years of industry experience.
> So anyone with claude access could arrive at the same conclusions ... and ask claude to fix it?

Begs the question of why the original "author" of the code hasn't just asked Claude to fix it? Or asked Claude to generate good code from the beginning. I suppose the answer is that nobody cares about good or efficient code anymore. But that's been the case since long before LLM coding though (as stated in your point 1).

Why write clever code when we can just write JS slop and ask customers "upgrade" their hardware every year...

It's worse than anything in MS really.
>2. So anyone with a claude subscription could arrive at the same conclusions?

It's symmetrical: just like anyone with a claude subscription could arrive at the same vibe slop!

Hmm the list is a bit underwhelming. Basically, it's unnecessary requests, bloated JS, unoptimized images and generally poorly structured code. I would hate if that was where the average website is headed, but realistically, we were already headed there before LLMs. From the headline I was expecting CVEs, broken UX flows / business logic, leaked secrets.
It's good that no leaked keys were found in the front end code but the developer wasn't able to look at the backend where, if the quality of the front end is any indication, there are likely to be many security issues. Hopefully it's not all running on the same servers/network as anything important.
That's because only frontend was reviewed, who knows what's on backend.
I love that they made the comparison to Hacker News in terms of request sizes lmao. Yea definitely a representative sample
And with that 37k a day, or 185k per 5-day work-week, what exactly has he achieved?

I don't see any actual output coming from these AI tools, despite how many are saying it's greatly increased their productivity. Where are all the new businesses and tools? I only see more shovels being sold.

Seems to be a frequently occurring issue, most (good) software engineers know that LOC basically means nothing, if anything, less LOC is a better goal over more LOC, if you absolutely have to have a goal. As soon as people/companies starts bragging about how many LOC they can ship, you need to start being very suspicious, mostly because they just admitted to not actually understand software engineering at all.

Cursor did something similar months ago, bragging about producing millions of LOC while what they actually made barely worked and could have been built with an order of magnitude less LOC: https://emsh.cat/en/one-human-one-agent-one-browser/ (https://news.ycombinator.com/item?id=46779522 | 324 points | 5 months ago | 156 comments)

What I don't understand, isn't there a single engineer working with these people who ask them what the fuck they're doing, before they hit that publish button? Or is there just such a constant pressure to publish anything that quality just simply doesn't matter at all to these people?

> The larger point is that while AI coding tools make it easy to pump out lots of code, it’s really (still) the quality of the code that matters. Quantity, in other words, doesn’t necessarily equal quality.

They didn't provide any evidence for this point. They showed that the code is bad quality, but not that it matters that it's bad.

How can anyone in their right mind think that growing a codebase by 200K LOC every week could possibly be a good thing?
Whilst I don't _really_ consider 37k LoC a day to be a particularly extreme number, my issue with all the high profile high output usage of AI, is that it doesn't really prove much of value.

I use AI in a pretty single-threaded way and my primary challenge is figuring out how to keep the LoC as minimal as possible, as well as minimise my spend. I am, unintentionally, one of the highest spenders at work; this is possible because I do more fiddly UI enhancements where the desired behaviour is often subjective and invariably hard to articulate. When I work on big backend features, my spend tends to be much lower.

What I really want from the people who have effectively unlimited tokens to spend, is to use that finding ways for the rest of us to produce higher quality output at lower costs, rather than focusing on output alone.

Nice article...

Bloat is not "since generative AI coding" new. We always had it. As the author says:

> “It does sound like Facebook’s ‘move fast and break things,’ which didn’t age well either.”

Which indeed did not age well, but it did help the company to grow to a certain point at which it is now a staple in our lives (and can do super expensive BS experiments like "Metaverse" and still show profits).

This may be what Tan does as well: first profitable, then correct.

This is an approach, which may work for some, it may also allow some companies to become irrelevant (like: once the bloat-app is profitable, a clone-app emerges that is bloat-free and overtakes the bloat-app in every dimension, while the bloat-app is figuring out how to scale up with a shitty db schema).

Do you remember when you had to write essays in school and the word count mattered? That’s when you’d write "it is" instead of "it's" and "will not" instead of "won't."

This is worse.

Not the first time that I read quantity over quality related to YC

But web dev is extremely bloat and inefficient already.
> Once Garry’s agent observed the usage data of his site, it would have corrected all these mistakes without Gregorein writing the thread.

That struck me most and that's why we (software developers) are never going to be engineer.

Consider a doctor just trying ang going along if the patient survives. A brige constructor doing something and after each car that passes fixing problems. A car manufacturer trying new brakes in production without testing. A builder starting to build a house and fix it while people already live there. The list could go on.

It's sad to see a profession on such a decline. I use AI myself, but please please keep doing something _professionally_ or stop doing it as your profession.

This is what happens when you give people tools that let them achieve an outcome, without necessarily giving them the judgement or expertise to know whether the outcome is any good.

If you asked me to build a house, I could probably assemble something that would stand for a few months. Hopefully. It might even keep the rain out. But it might also fall on my head, because I do not know enough about building houses to be confident that it won’t.

And even if it didn’t fall on my head under normal conditions, I also would not know when I needed to design for earthquakes. Or floods. Or fire. Or wind. Or grandmother-cosplaying wolves with very strong lungs.

But if all I need is shelter for a day, would I necessarily care whether it lasts more than a week?

That is effectively what a website like this is. It is not really a product. People don’t depend on it. Tan’s visitors are probably using MacBooks and iPhones on fast networks, and most of them will never notice how bad it is under the surface.

That does not mean it is good. It means it is good enough for the context.

Most people also tolerated the hilarious gigabyte JSON parsing bug in Grand Theft Auto for years, until a hacker patched it and cut GTA Online loading times by around 70%: https://nee.lv/2021/02/28/How-I-cut-GTA-Online-loading-times...

It was good enough, even if people noticed how bad it was.

Business applications, and typical software really doesn’t have to be super tuned or perform fast. It just needs to work.

At least until your product category has been commoditized, and _then_ you’re competing on experience.

You can make decent code with AI assistance. You can churn out 37kLOC per day.

However, I’ve seen no evidence to suggest that you can do both.

> Tan/AI built the website so that when a user visits, their browser makes 169 server requests for various assets totaling 6.42 megabytes in size.

No, no, that's just average day of frontend development

Garry is buying horses for all the defenders in this thread.

No he wont fund your startup.

That's just sick. If I had to review 37k LoC per day, my brain would vaporize. But the point probably is that nobody ever looks at that code.
Focus on a single metric instead of outcome and you win on that metric instead of the outcome.

I remember that for, uh, Key Quarterly Objectives, was that the name? Aeons ago.

Same shit new decade.

(I love working with AI though. It has many of the benefits of good pair programming.)

Does anyone know which CDN this site is using? I got an unusual type of CAPTCHA asking me to slide right.
Garry Tan's point still stands: he never pretended to be building nice software. But his point was that he can now build AT ALL! Shipping a webpage at all is the firs step ; making it load under 7 Mo is just a refinement, an important one of course (who tf wants bloated webpages) but still only a refinement. Tan is right to be amazed and to be shipping.
The cult of 10x got their magical tool that lets them feel 100x. Don't take their toy away from an excited kid.
Are there ANY institutions remaining that haven't completely succumbed to the Church of Quantity? It was already bad before AI, on the web especially, and now that it's widely used it feels like no one cares about anything anymore. Ask any developer and they would gladly offer you 100x the code shipped 10x faster, but recoil at the idea of making it even 1.5x better or more performant. Everyone has accepted that quality doesn't matter because... you can remake the same flawed thing ten times in an hour? Code is discarded just as quickly as it was spawned into existence, never having been worth anything to anyone. But at least some very influential people get to take the speed they firehosed this code at and brag about it. Not what they made, not its quality, but about a metric that only builds street cred among a million people doing the same thing and not really going anywhere.

I knew that the elite of the tech world never really cared, they were just forced to. They don't care about what they make as long as some numbers and statements look good so they get more money at any cost. What's shocking is that everyone else agrees with them - that in all contexts, quality is dead and less than worthless. Just stop caring bro.

An opinionated agent harness is all it takes to correct the mentioned issues.
"Tan/AI built the website so that when a user visits, their browser makes 169 server requests for various assets totaling 6.42 megabytes in size. For comparison, the minimalist Hacker News homepage (also run by Y Combinator) makes seven requests for data totaling just 12 kilobytes.

The website ships 28 actual test files (code developers use to reality-check their work) straight to every visitor’s browser. That’s 300 kilobytes of pure developer scaffolding that users never asked for.

It loads 78 different JavaScript controllers for features like AI image generation, voice extraction, video tools, etc., none of which appear on the homepage. The browser still has to download all of them “just in case.”

The site’s logo is an illustration of a bear. The site downloads the logo in eight different formats, including a completely empty 0-byte file that somehow made it to production, Gregorein found.

The website uses huge, uncompressed old-school PNGs (some nearly 2 megabytes each), even though the browser literally asks for modern tiny formats. Two images alone waste about 4 MB; with newer formats they could have been just 300KB.

Gregorein also found duplicate page content, an empty CSS (Cascading Style Sheets) file, a huge rich-text editor loaded on a read-only page, missing image descriptions, and analytics code that deliberately routes through a proxy to dodge people’s ad blockers (with a comment in the code admitting it), Gregorein reports,"

Brave new world indeed... It might be elliptically relevant that in the original novel by Aldous Huxley, the protagonist John the Savage hangs himself at the end when his search for truth fails.

my cpu fan spinning a lot when visiting that website
> The larger point is that while AI coding tools make it easy to pump out lots of code, it’s really (still) the quality of the code that matters

This conclusion is unsupported from the observations. The code makes lots of requests, has too much CSS, and 6 different logo formats. So what? Lots of real, production codebases have just as many warts.

Folks need to stop dealing in absolutes with AI coding. Code quality always mattered, and still does, in certain circumstances. In others, it's less important and speed of iteration has more value. That's still true even with AI code.

It's 100% fine for Gary to write software in a way that works for him - and for us to push boundaries, and not 'gatekeep'.

BUT - it's still a pile of steaming garbage.

It's not just a 'bit odd' - it's just massive slop.

The total lack of self awareness is comical and disturbing.

GStack + Gary's Tweets about 30K LOC a day is 'Trumpian' in self delusion - it looks like basically he and YC are frauds that don't know what technology is.

This is like Elon and his 'salute' or his 'I'm a top 10 Diablo player - hey watch me play and expose myself unwittingly' type situations.

A bit of AI hype is fine.

Someone needs to take Gary and have a side-discussion.

He can do GStack - he just needs to characterize it properl for what it is.

Talking about that amount of code with the assumption 'it's sensible' code, is basically a lie - it's fraudulent' levels of hype.

Just describe it as an unwieldy but productive experiment, not for mainline consumption etc. then it's fine.

Ironically the fastcompany page hosting the article that should only be read only paragraphs of text and a simple navigation seems to do 300+ requests, with 29.4mb of resources.

I mean honestly the numbers and issues talked about there seem to be on most modern web apps. It just doesn't seem to affect business results enough for people to care about fixing them.