back

by dhorthy·10mo ago·view on hn ↗
if you read further down, I acknowledge this

> While the cancelation PR required a little more love to take things over the line, we got incredible progress in just a day.

2 comments
FWIW I think your style is better and more honest than most advocates. But I'd really love to see some examples of things that completely failed. Because there have to be some, right? But you hardly ever see an article from an AI advocate about something that failed, nor from an AI skeptic about something that succeeded. Yet I think these would be the types of things that people would truly learn from. But maybe it's not in anyone's financial interest to cross borders like that, for those who are heavily vested in the ecosystem.
But, yeah, looking again, that was a pretty big omission. And even moreso, a missed opportunity! I think if this had been called out more explicitly, then rather than arguing whether this is a realistic workflow or not, we'd be seeing more thoughtful conversation about how to fix the remaining problems.

I don't mean to sound discouraging. Keep up the good work!

there is a portion in the article where I talk about how our hadoop refactor completely failed
I think what the OP is asking for is an article _like this one_ about where you go in-depth into what you tried, where the system went, and more specifically what went wrong (even if it's just a list of "undifferentiated issues"). Because "we tried a thing. It didn't work. We bailed out." doesn't show off the rough edges of the tool in a way that helps people understand "the shape of the elephant".

Or, in the vein of https://adamdrake.com/command-line-tools-can-be-235x-faster-... - "here's a place I wouldn't use an AI tool because _other thing_ is far better"

You do acknowledge this but this doesn't make the "spent 7 hours and shipped 35k LOC" claim factually correct or true. It sure sounds good but it's disingenuous, because shipping != making progress. Shipping code means deploying it to the end users.
I'm always amazed when I seen xKLOC metrics being thrown out like it matters somehow. The bar has always been shipped code. If it's not being used, it's merely a playground or learning exercise.
would "wrote" be more appropriate than "shipped"?
the most accurate wording here is "generated".

we generated 35K LOC in 7 hours, 7 days of fixes and we shipped it.

This at least makes it clearer that it is on par with what it would take a senior BAML team member to accomplish this, which is kind of impressive on it's own. Not sure about ignoring the tests though

I don’t think it is any better. 35kLOC of slop isn’t a good metric by any measure, no matter what word you use.