Hacker News

Favorites Setup
Comment by chamomeal | original | Grieving the loss of details
[−]chamomeal · 2026-10-10 Sat 21:02 UTC · link
I agree with many perspectives in this thread. I empathize with OP. I’m scared about the future of my job and already feel it less fulfilling, despite being more productive than ever.

But I have one singular counterexample to most scary narratives and essays about LLMs writing software, which is my coworker who hardly uses LLMs at all. At least, he hardly uses them compared to me. He still writes most of his code by hand. He still greps around the codebase without claude code. Like he’s definitely using claude code, but only for like one-off specific tasks.

He’s a totally average developer, like everybody else on my team (including me). But he’s clearly more helpful than the rest of us. When people from other teams have questions about how something works, he’s always the first to respond. He’s always the one providing useful context in planning calls. He catches stuff in code review that I didn’t catch, and claude/copilot didn’t catch.

I think I’ve leaned too far into AI, partly because I’m a lazybones and am kind of burned out already. But there’s a stark contrast between me and my coworker that there didn’t used to be. I think the context rot is really setting in, and getting worse. So I think there’s still value in caring about the details

[−]enraged_camel · 2026-10-10 Sat 21:37 UTC · link
>> He catches stuff in code review that I didn’t catch, and claude/copilot didn’t catch.

This, I find completely unbelievable. Because we had (emphasis on had) seasoned engineers on the team who similarly eschewed AI tools and insisted on doing everything by hand in the manner you describe, well after the rest of the team adopted AI. Yet when it came to code reviews, even in parts of the codebase they were familiar with, the bugs they found often came down to nits, bike-shedding and opinion-based feedback (that they usually tried to frame as objective fact). They would also disagree with almost every AI finding, arguing that it was an unrealistic scenario or an edge case not worth worrying about.

Fundamentally, I don't think humans are going to be capable of providing high quality feedback on PRs authored by AI agents unless those PRs are fairly small in lines of code and volume. It's just way too much information and context for one person to keep in their head. I read a statistic that said the average lines of code a senior engineer can read and provide good feedback on is about 400 per hour, and that number goes down the more time they spend doing code reviews. So, to anyone who insists on trying to keep up with AI, I say: good luck.

[−]geraneum · 2026-10-10 Sat 21:55 UTC · link
> This, I find completely unbelievable.

You both have anecdotes. Anecdotes don’t “cancel” each other out.

Here’s a third one. In some of the code reviews I’ve encountered that AI gives a lot of feedback, it’s just providing noise. Things that should be ignored or when following the feedback causes more harm which requires more token to “fix” later on. That can also happen. Sometimes the thing it spits out goes against the common sense, and sometimes it works very well.

> So, to anyone who insists on trying to keep up with AI, I say: good luck.

This I agree with, for a different reason. It’s like trying to swim in a sea of honey and trash mix. It’s exhausting.

[−]abuani · 2026-10-10 Sat 22:14 UTC · link
> Here’s a third one. In some of the code reviews I’ve encountered that AI gives a lot of feedback, it’s just providing noise

Something I've found fun is seeing how long it takes for an llm review tool to come back satisfied with a PR. Think 100 lines of code changed, nothing terribly significant, but also not trivial. I'll have a local Claude session setup to babysit the PR and wait for feedback, accept all the recommendations, push the change up and request a review. I cap the number of iterations at 10 just so I'm not blowing a stupid amount of money. I've yet to come up with a PR where the llm reviewer is satisfied with the changes and has _no feedback_.

So where's the reasonable cutoff point for llm based reviews?

[−]bgoated01 · 2026-10-11 Sun 00:20 UTC · link
Huh, I've had many times when `codex /review` comes back satisfied on the first shot, both with handwritten and LLM-assisted PRs.

We use only one round of LLM review, and have discussed as a team still assessing recommendations, not just accepting everything blindly. So somewhere in the (0,1] rounds of review. Sounds like our preferred ratio of human to LLM involvement is different than yours, though.

[−]enraged_camel · 2026-10-10 Sat 22:32 UTC · link
>> Here’s a third one. In some of the code reviews I’ve encountered that AI gives a lot of feedback, it’s just providing noise.

This is completely normal if you haven't written your own custom skill with instructions on what types of issues the AI should emphasize, which ones would be considered nits and which ones aren't a problem at all.

In our repos we use a classification system: blocker, should-fix and nit. Each one has specific definitions, criteria and examples encoded in the skill file. When the time comes to review a PR, agents invoke the skill, and frankly do a stellar job. A human then reads each finding, asks the AI follow-up questions and makes the final decision in terms of whether the finding goes in the PR review.

The reason I know this works is that we have one guy on the team who does not use this skill, and blindly throws his agents at PRs. And the results are exactly as you describe.

[−]LoganDark · 2026-10-10 Sat 23:49 UTC · link
> This is completely normal if you haven't written your own custom skill with instructions on what types of issues the AI should emphasize, which ones would be considered nits and which ones aren't a problem at all.

"You're holding it wrong" :)

[−]lmz · 2026-10-11 Sun 04:42 UTC · link
> "You're holding it wrong" :)

It's a general purpose tool. There are different ways of applying it, with different results.

[−]geraneum · 2026-10-11 Sun 05:05 UTC · link
> This is completely normal if you haven't written your own custom skill with instructions

This line of thinking comes from the assumption that LLM are somehow infallible and can’t be wrong, which is obviously not the case.

Writing skills and other magic incantations is the first thing that comes to mind of anyone who sees the review results for the first few times. What makes you think we didn’t do that?

[−]majormajor · 2026-10-10 Sat 22:24 UTC · link
Claude's code review skill, in particular, can find some good stuff. But it has some big blind spots around certain types of code. And it likes to come up with a lot of nits too—I think it's really really trained to try to always find between 2 and 8 things or somesuch. Good news is that it is very receptive to "nah" on the bikeshed ones and doesn't stick with them, but will stick with big issues. It'll probably bring up a few more nits though that it didn't bring up the first time!

But I can completely believe that someone who knows the code by heart would have a better signal to noise ratio on their reviews.

I'm trying to find the sweet spot because I've found some NASTY bugs Claude missed, and also had Claude find some nasty ones for me. And this is in codebases with tens-of-thousands of AI-generated lines of code + AI-driven reviews. So I want to bring both to the table.

The existence of some of these major "oh man that changes a lot of our assumptions" bugs that were only found because someone poked on the agent and said "I don't think you're paying enough attention to this" justifies that, IME.

And the better you are at pointing the agent at the truly-important parts, the better the agent's gonna be at finding shit you missed.

[−]XorNot · 2026-10-10 Sat 23:12 UTC · link
Claude's code review is a lot less interesting then getting Claude to reproduce the bugs it claims to find in code review, which has had an absurdly high hit rate for me.

The biggest problem I see with how a bunch of people use these tools is they go to them as an oracle, rather then letting them be plugged into and interactive with problem.

And it's in that later context that Claude is amazing: it can run tests and setup scenarios which would take days or get stuck in some weird problem loop. And then you can just say "okay, walk me through this problem" and see it yourself right there.

[−]palmotea · 2026-10-11 Sun 07:40 UTC · link
> The biggest problem I see with how a bunch of people use these tools is they go to them as an oracle, rather then letting them be plugged into and interactive with problem.

At work they have some scale of how "advanced" of an "AI engineer" you are.

IIRC, using the model interactively means you're stuck and "level 3." IIRC, level 5 (the best) is having some agent interview you about what to do, generate a story from that, then some other agent consumes the story and implements it, etc. I think you're supposed to check their work at each step, but that sounds inhuman and unfulfilling.

[−]jltsiren · 2026-10-10 Sat 23:31 UTC · link
A lot of bugs are essentially that the code does something plausible correctly, but it's the wrong thing to do. People who are familiar with the codebase and its purpose can often spot those bugs. Those with less experience usually can't.

The issue is tacit knowledge, or implicit context. Most details are never written down. If you don't know them from experience, you have to guess. If you guess, you often guess wrong. And even if you know something from experience, you are often not aware of it, until you see something that violates it. So it's not possible for you to write it down in advance.

You can replace people with AI in the above, and nothing fundamentally changes.

[−]StrangeWill · 2026-10-10 Sat 23:40 UTC · link
As someone who owns a company:

Some holdouts excel at what they do to the point of retaining significant value, earnestly worried about being out of touch with what we do, some holdouts just suck and are bitter.

[−]zzzeek · 2026-10-10 Sat 23:23 UTC · link
I deal with a lot of non-LLM users at work and quite often the conversation hits a dead end, where I'm like , ok but see, your app has this performance blockage due to your not having an index here - do some performance tests and you'll see (it was a hang caused by a FOR UPDATE locking the entire table due to lack of an index). And you can hear the pause (it's all over Slack) where they just aren't going there, because writing a performance suite for the issue in question would be a lot of effort to do by hand and in the "before times" would be a difficult undertaking to justify. Because they don't consider an LLM, a task they most certainly should be doing becomes a non starter. Never mind my own Claude had a whole plan ready to go to do this whole suite for them in about five minutes so they could study the impact of the index, but I really didn't want to just go ahead and do this all for them. At some point your non-LLM coworkers need to stare directly at work we'd never be willing to do before that's now trivial, and in fact is now part of the job.
[−]nostrebored · 2026-10-11 Sun 00:07 UTC · link
Right, but performance testing has always been hard and LLMs are awful at it. So maybe your coworker was right?

I’m definitely on the “just use the clanker” end of the spectrum. But knowing when and is good. Constantly having to argue with people’s reposting of Claude’s idea of what’s wrong has left me burned out. At this point, if I have any idea about what’s going on and someone tells me “Claude says”, I will ignore them.

In this case, it sounds like you could have made a ten line repro that shows the problem. Why not just send that, something human interpretable?

[−]cracell · 2026-10-11 Sun 01:41 UTC · link
What LLMs are you using? Just simply telling Sol 6.1, Opus 5.5, Fable or Astra to performance test something will get you a solid improvement in poorly optimized code.

Giving it a specific plan will get you a solid test harness.

And setting up an autoresearch system and running it overnight will get you expert level optimizations if you set the metric up right.

These things excel at performance optimizations.

[−]nostrebored · 2026-10-11 Sun 04:43 UTC · link
Using everything modern and useful!

Performance testing has always had problems with isolation, mocking, covariance of services etc. I’ve just spent two weeks driving down latency across our framework, and autoresearch was definitely not a viable path. Most of these loops have these logarithmic, non-step change curves.

Just like with actual research, the real benefits came from me asking why is it like X and not like Y?

The original performance testing framework Claude created to bench our different versions against did not even mock high variance provider calls…

[−]slopinthebag · 2026-10-11 Sun 08:28 UTC · link
yes but you can use an llm to do the grunt work which still speeds up the process, you don't need to be like "computah speed this up", you can be like "computah, write me a profiling script using lldb to inspect this one hotpath and look for X, Y, and X" etc etc.
[−]zzzeek · 2026-10-11 Sun 04:05 UTC · link
> Right, but performance testing has always been hard and LLMs are awful at it. So maybe your coworker was right?

LOL, you're making the same mistake they did. thinking "index" means "too much time spent fetching the rows". read again - it's a FOR UPDATE so the entire table gets locked and other processes get totally blocked.

> Constantly having to argue with people’s reposting of Claude’s idea of what’s wrong has left me burned out.

It wasn't just claude's idea, it was my idea too, the claude topic was that it would write a suite that proves the problem, in this case, very loud logging messages that were occurring for the customer when this quasi-deadlock situation occurred. it was not subtle.

> In this case, it sounds like you could have made a ten line repro that shows the problem.

no, it involved running a galera server and about four other services with a specific set of data conditions, again, read what I wrote, creating a proof of concept suite was not trivial to do by hand.

[−]nostrebored · 2026-10-11 Sun 04:48 UTC · link
But why would you have to make the succinct suite by hand? Not what I was suggesting, I have gotten in the habit of sending extremely concise, human readable repro scripts made by Claude.

Partially because I can read and validate it actually shows what it claims to show, and partially because I expect people want to know what I’m telling them.

[−]sgarland · 2026-10-11 Sun 12:38 UTC · link
> it's a FOR UPDATE so the entire table gets locked and other processes get totally blocked.

You’d think they’d have found this issue out before now; if there’s no usable index for the predicate, non-locking reads would’ve also been slow.

Also, I look forward to the next update wherein your colleagues create an index, and then discover the joys of gap locking under REPEATABLE-READ.

[−]dev1ycan · 2026-10-11 Sun 02:02 UTC · link
This is me basically, I use LLMs on the browser for specific queries rather than you know, have it abstract my work into a black box...
[−]abalashov · 2026-10-11 Sun 04:14 UTC · link
> I think I’ve leaned too far into AI, partly because I’m a lazybones and am kind of burned out already.

I've been doing systems / backend programming since I was 10 (in C in those days), so I was burned out by the time I was 20, and I'm now 40. The temptation to trade on that knowledge in a very short-term way is strong, but I know where it leads -- to the rot you describe.

[−]xtajv · 2026-10-11 Sun 07:40 UTC · link
This is rotten but I'm secretly glad that Claude et al are metering by the token.

I'm not sure that folks realize that the purpose of software engineering is to program in a way that doesn't just function, but optimizes for future readability, maintainability, and behavioral change in line with product requirements.

There's a reason why programmers are not paid "by the line", and it's because LoC as an incentive structure is a disaster that leads to brittle and verbose slop code -- and we knew that even before LLMs.

So I'm glad that the (long-term) economics mediate against token-hogging software by making it literally more expensive to deal with.

I just hope that managerial types are smart enough to realize that "time to understand/upgrade/change a codebase" still matters, whether that's expressed in terms of SWE work hours or LLM tokens.

[−]abalashov · 2026-10-11 Sun 07:47 UTC · link
I would not bet on the managerial types responding to anything but (a) blunt monetary incentives and (b) the latest buzz from other people in their general professional category, e.g. at the country club.
[−]spacechild1 · 2026-10-11 Sun 11:58 UTC · link
This is one (of several) reason why many larger open source project ban or severely restrict the usage of LLMs. They want contributors to fully understand the changes they are making.

Recently we have received several bug reports with AI suggested fixes. They looked reasonable enough and seemed to resolve the issue at hand. I then took a deeper look and discovered that every single suggestion only hid the symptom but did not address the underlying issue. Applying the suggested fixes would have hurt our codebase, so I eventually had to reject all of them and write my own solutions. This was only possible because I'm deeply familiar with the codebase and the broader context. Someone with less knowledge would have probably just applied the LLM suggestions without much questioning. I find this rather worrying.

I guess in popular open source projects the stakes are just higher than in some proprietory software that will be replaced in a few years anyway. The irony, of course, is that LLMs have mostly been trained on open source code.

Also, some of the projects I'm involved in are over 20 years old. One even celebrated it's 30th anniversary. This means there is also the question of historical responsibility.