back
249 comments
>On my visits to the Bay Area, I would ask AI researchers or interns why they are doing their current research or projects, when in a year or three agentic LLMs could probably do them;

This is such a weird point to make that doesn't become correct just because everyone makes it, all the time. Why clean the ocean if some magic future tech will clean them? Why save the world now if some benevolent AI is 'just around the corner' and will do it for us? And people have been making this point for years now, and it's not like my job got any easier. I just got more AI.

https://www.poetryfoundation.org/poems/51294/waiting-for-the...

And I say that as someone who uses Claude Code in complex environments almost hourly; I, as the human, still have to do the thinking as Claude still 'can't jump' [1] and I have seen no evidence that they (or similar AI, any time soon) will 'jump' like a human brain does.

[1] https://www.tomzahavy.com/files/llms-cant-jump.pdf

> This is such a weird point to make

I think it is a great point to make, because if everyone really believed that AIs will do everything without human intervention in a handful of years, as the marketing repeats again and again (AGI, singularity, etc.) and have been saying for years... why then get bothered?

Because we DO know LLMs have their hallucinations, limitations, perform tasks not previously seen way worse than humans, etc. And it seems that, for now, there is not a good or magic solution to it, it is inherent limitations of the paradigm.

Yes, you can feed more and more and more (curated data) and eventually make AIs excellent at task X or Y, but then you spend your time specializing those engines. So the work does not really disappear, it just shifts and you make it more replicable for a bound set of problems.

Needless to say that at some point I prefer to learn (and combine with AIs, it is ok) than acritically getting inputs from something until I become totally useless.

Unless we have a paradigm for which a fully autonomous AI can do everything, this will just become improving our productivity in some ways, with all the in-between bottlenecks that it has.

I think you meant to post this in response to https://news.ycombinator.com/item?id=49174900?
> I, as the human, still have to do the thinking as Claude still 'can't jump'

I still have to do quite a bit of thinking but the amount of of thinking I do per task is trending down. I agree LLMs are not good at abduction but very few humans are either and very few jobs/tasks require it. I can't talk for researchers jobs though. But perhaps fewer researchers would be desired by these labs (not none).

Same reason some think preserving the environment is pointless because the believers will ascend to heaven, either way. It’s a religion. It’s dogmatic nihilism.
By trade I'm a UX Researcher/Designer who designs in code (HTML/CSS) and have done so since 2009. Recently I vibe coded an entire python app with a database and each time I didnt know what to do I would just feed screenshots to Gemini or Codex for guidance (i think i could share my screen with Codex and it can guide me via a voice conversation). I know I could follow up and build a companion iPhone and Android app using these tools.

Overall, I'd like to understand those who have a positive outlook on design and software engineering as a career. Where do you see the opportunity where I just see a bleak one where anyone can do this stuff by typing or talking to AI? Myself, after 17 years in the field I am begrudingly back in school for a new medical career. As well, anytime an IT recruiter reaches out I am getting responses back only after under-cutting the hourly rate I use to demand and what others probably are still trying to get. And with it feels even bleaker as it becomes a race to the bottom!

It reminds me of Richard Hamming’s notorious question. I like the summary at https://bestjelly.substack.com/p/hamming-questions (which starts out with a quote from another site):

> > Mathematician Richard Hamming used to ask scientists in other fields "What are the most important problems in your field?" partly so he could troll them by asking "Why aren't you working on them?" and partly because getting asked this question is really useful for focusing people's attention on what matters.

> I imagine someone being asked this question, and how they should respond. I think like so - ‘Fuck off Richard’.

> This is partly because I imagine this question being asked in a kind of snarky, gotcha kind of way, with some sort of nerdy superiority. Like ‘ha your behaviour is inconsistent with your implied preferences, you idiot, do you even von Neumann–Morgenstern?’

although, if i'm out of tokens and have to wait a full day, i won't bother doing some things manually because the day i'll spend doing something won't take more than 1 hour the next day when tokens are available again.
Also people tend to forget that LLMs still just work on compressed data... Where are the MAJOR breakthroughs? Where is all the "crazy" AI output going? Software seemed to degrade in quality a lot in the recent years. All "improvements" LLMs go through are simply improvements on how to burn more tokens out of my pockets given that Claude now want an actual browser extension to "visually" confirm small changes every time I use it for UI. They are still just data parrots.
> And I say that as someone who uses Claude Code in complex environments almost hourly; I, as the human, still have to do the thinking as Claude still 'can't jump' [1] and I have seen no evidence that they (or similar AI, any time soon) will 'jump' like a human brain does.

Sure it can, turn up the "temperature" a bit.

There's this notion that human "jumping" is magic. It's not. It's all based on inputs. Including unrelated inputs, past inputs, and feeding yourself your own thoughts.

The hard part is not the ability to make conceptual jumps. That's just random search. The hard part is discrimination: whether a given mental jump is "creative" or "insane". Iterated, the problem is that of balancing between the two failure modes: relax your thinking too much, and you'll start thinking nonsense thoughts; tighten it too much, and you'll be just following immediate-term rewards and obvious thought trains. It takes time to find that balance, and plenty of people at various points err in one or the other direction (e.g. small kids in particular tend to err on the "crazy non-sequitur side", but that's because they're learning the basics of reality and social interactions).

> We already know developers don’t actually spend most of their time writing code, with studies at Microsoft and elsewhere showing it’s closer to 14 percent.

Anyone else finding they're spending more time writing code (or at least driving agents to write code) now?

14% used to feel about right for me - I'd spend the rest of the time researching approaches and libraries, planning things out in issues, or sometimes just thinking really hard about problems I ran into.

Now... I still do those things, but I'm doing many of them faster - and I'm often doing them while my coding agents are churning away on code.

There's also this weird effect where the harder a problem is the more I can get done in parallel with it, because an agent might need to spend 20 minutes on it without my involvement.

I've found that using LLMs for significant amounts of code generation completely drain the result from any dopamine I would get doing it myself.

Have others noticed this as well? This is going so far as to me losing interest in side projects because I have "lost touch" with the code base.

I feel like all you need to know about how seriously to take this is that they cite that ancient early-2025 METR study, and describe it in the text as "recently one even found..."
Like many others in the comments, I feel there are a lot of assumptions in this piece. Before, coding is only 14% therefore, small slice. I think that's a very superficial assumption. That was because coding was expensive and we needed to be sure we didn't code the wrong thing. If code is as cheap as it is now, we will optimize differently, we will structure around it. Instead of so many meetings we will code 5 different versions of the same thing and choose, etc.
> Myth 2: Writing Code Is the Bottleneck

Writing code is indeed the bottleneck for same resource constrained companies.

Rapid code development creates more opportunities for trial and error, providing companies with more information for decision making, that previously might have been addressed by meetings.

Of course, this might bring other problems, but it might not right to generally speaking that writing code is not a bottleneck.

This reads like a critique of 2023 tooling published in 2026. Their Amdahl-style arithmetic (speed up a 14% slice, cap your gains at 14%) holds only if "AI" means autocomplete. Current frontier models do far more than that: research, code comprehension, review, test authoring, debugging, exploratory prototyping, ideation. That's most of the rest of the working day or "86%".

The only point that still holds is that organizational policies and procedures that automate AI use and lower the barrier to entry are more efficient than leaving it up to each individual. Every other point they make is either stale or was never true to begin with.

I don't understand Myth 1 (Developers Spend Most of Their Time Writing Code).

They quote a study in which developers report to spend 11-14% of their day coding. The rest is stuff like solution design and meetings. The insinuation is that AI can at most automate 14% of your day.

The problem with this argument is that once you have code, some (not all) of the precursors to code go away.

|--------|-------|------|------|-------|------|

|Contract|Product|Design|Coding|Testing|Deploy|

Writing Code Isn't the Bottleneck, until writing code is the bottleneck, until it's not again.

I can't remember of any product in my lifetime that was more over hyped than AI.

It's a huge piece of shit and if I wasn't forced to use it at work I would never use it.

It writes dumb, throw-away code and adds thousands of dollars/developer in costs. All this crazy code that we're adding to our projects will come back to bite us in the future, there's no way it won't.

> f developers spend only about 15 percent of their time typing in the editor

I think this is missing an important detail. Lots of time was spent on non-coding stuff, because coding used to be more committal and hence expensive. With how quickly one can code up a quick prototype or even production-ready code these days, the code becomes the communication tool as well.

> Myth 3: Lines of Code Written by AI...

how come lines of code (or expressions) by an engineer aren't a good way to measure progress (Gates point etc) but GenAI tokens must count and be paid for?

> A study of more than 450 engineers at Microsoft in 2025 showed developers spend only 14 percent of their time writing code

This feels about right as an average across project cycles and different types of companies, and is the same type of number I've suggested here before.

The obvious conclusion is that even if AI reduced coding time to zero, then it would only reduce software development time by that 14%.

Of course AI may be used for other aspects of the job as well as coding, but on the flip side any serious use of AI requires a human in the loop to give work to the AI, steer the AI, assess the output, etc.

It's interesting that this 14% figure is close to Uber's choice to limit AI spend to 10% of developers salary. I wonder where that Uber number came from?

Another rather startling datapoint on the perceived value of AI comes from Microsoft who are also adopting budgeted AI usage, and are looking for "outcomes that move the needle". Their budget guidelines apparently refer to (current? targeted?) per-developer AI usage of "hundreds of dollars a month to a few thousand dollars in tokens".

https://www.techradar.com/pro/tokenmaxxing-is-not-what-we-ar...

It seems that this could have been expanded or contracted to any Fibonacci number of myths.
All very sensible points which I think all senior programmers who have used AI would largely to agree with.

For those more junior - keep in mind that a lot of the maximalist rhetoric are from people either selling models, or the cottage industry of people selling you courses or tools to help you use the models. Try and keep in mind software is not a mature industry, it's an immature one, and it's prone to hype and fads.

>>>Recent research shows that developers—especially women and older engineers—face a “competence penalty” when using AI

and why did this study (performed in China I might add) find that women and “mature age” (which is not defined in the study) have less usage of AI tools?

>>>? We suggest a new barrier: using technology to assist task completion signals a lack of competence to perform the task independently.

So basically the study suggests that older engineers and women are reluctant (and especially women) to use AI tooling because they feel they are being judged more on non-technical competencies.

This study may have a strong cultural influence but I would say that one thing they noted I’ve also seen. 41% of engineers in the study had used AI 12 months after the initial rollout. This aligns with my observations. Some people are struggling figuring out how to adopt in their day to day while others are full in.

Some experience from my work:

- In biz development, a dev usually spends 30-40% time on coding, and more time on requirement discussion, integration testing (especially when the tests involves mobilephone or car)

- coding time can be reduced to 30%, which means reduce 20%-30% time of the full pipeline

- meanwhile, every phase and role is using LLM now, for example, product manager can produce longer requirement doc easily (we can use LLM to read it anyway:) Meeting sometimes is more than before, because more document output leads to more reading and discussion.

- I hope to find new ways to express biz requirements, in a more efficient and automatic manner.

- Shorten the requirement-dev-test-deploy loop is very important. OUTPUT is not OUTCOME. It is equal when we can see the final result, instead of intermediate metric.

- Agentic infra is extremely useful, or every one will find a way to access the database, report and ops system, in some weird fragile method.

I think the paper would have been stronger if it acknowledged how quickly the underlying evidence is becoming outdated. AI-assisted development in 2026 isn't just better models. The way many devs including myself work has changed and matured quite a bit as compared to last year
This is actually true at my company. They expect employees to be 10× more productive now that we have AI.
11-18% of time spent in coding is still very high number I think. For a large org with lots of process and risk aversion, this number could be as low as 5%. Even for 14%, the 10x improvement could mean 86+(14/10) => 87.4/100 => 12.6% overall time saved.
>a “good” workday, engineers spent 18 percent of their time “coding” (not including bug fixing, testing, etc.)

that's such a weird metric, why exclude bug fixing and testing? depending on the phase of the project I might spend 100% of my coding time bug fixing

I look askance at an article that uses a GenAI survey from 2024. Other references could be good. But the dev agent milestone was Nov/Dev 2025, and CoPilot and agents were pretty sketch before then (but helpful at times.)
A lot of this rings true, but I think it's still too narrow. Sure, coding does not equal productivity, that is well debunked already. But I would argue that productivity is a product of engineering delivery + product decision making. Now where is the line between product and engineering? It varies by company, team and individual, but I don't think productivity can be measured for those functions independently, and in fact I see gains from AI on both the coding AND the product management side.

Basically as a senior tech lead in a large company engineering org, I don't have the bandwidth to individually validate every assertion from engineers on other teams OR from every product manager that comes with a half-baked ask. In the past I would be limited by the influence I could get through human relationships to strong SMEs with good judgment, and those folks always thin out as a company grows and calcifies. The number of creative and innovative thinkers dwindles, and the number of people protecting their turf and doing the minimum not to get fired increases. As a result many good ideas can get blocked by random gatekeeprs with poor imagination, poor expertise or both. However with AI I can follow up on gut instincts and fact check a lot more things, and ask incisive questions that can cut through a lot of organizational bullshit.

That's where I think most of the AI gains are today. Of course once AI plateaus and normalizes I think it will be baked into the org structures of tomorrow. But for now it offers real competitive advantage to those with the expertise to ask the right questions.

> studies at Microsoft and elsewhere showing it’s closer to 14 percent

This is a depressing stat. The real productivity gains come from leaving soul sucking big tech companies where nothing gets done with any sort of urgency.

If you can spin up a new software project with little to no effort, the most important part of the job becomes making the right decisions and this is where we will be spending 99.9% of your time.
With all the myths and hyperbolae circulating regarding AI, I'd love to know what it's like at large software companies adjusting to this brave new world.

It's easy for a small team to adjust workflows and roles, but I just imagine the office politics must be a waking nightmare in big organisations right now.

> Despite this decade-old research, many organizations still rely on lines of code as a measure of developer productivity.

No, they don't! It's easy to dispel myths when the myths are built on straw men. Dumb article.

I mostly agree with the part where it is stated that AI is a tool which received massive investments without knowing how to maximize its utility. I think that we will pay for it in the near future.
>a “good” workday, engineers spent 18 percent of their time “coding” (not including bug fixing, testing, etc.)

I must be a crap developer, because I probably spend twice as much time bugfixing and testing than "coding". (Both of which actually involve coding stuff, so I really don't like that distinction they make)

This is stuff AI can be really good at, so brushing that part under the table distorts the picture.

Having said that, I do agree with most of the myths they present.

It's kind of weird how we blame agents for hallucinations as if humans don't fall for that as well, while agents can run the build-fail-fix-repeat in much faster cycles than coders.
I hope that some of the executives out there will read the list. I know they won't spend the time to read the full blog, but at least the headers should be enough
> A June 2025 study of Microsoft developers

A year ago feels like forever

> AI Will Turn Individual Developers into 10x Developers

Maybe not 10x, but there's definitely a significant boost in velocity of development and delivery.

We used to be a team of 20+, now only 9, and we are delivering more than before. Besides that, we did some major refactors that were sitting in the backlog for months.

one still has to think. Also about this ai automating stuff and humans playing around. I believe there is a time for this and time for that
even an AI assist that makes coding twice as fast would, in theory, improve developers’ overall productivity by less than 15 percent. The other 85 percent of their time remains untouched

I stopped reading after this. AI has massively impacted most aspects of my non-coding work including the mentioned planning, understanding legacy code bases, setting up environments, etc etc.

Either this article is written by people with skill issues or - given the platform - its a biased and protectionist take that will fall quickly under the march of reality.

ChatGPT, summarize