Hacker News

Favorites Setup
Comment by geraneum | original | Grieving the loss of details
[−]geraneum · 2026-10-10 Sat 21:55 UTC · link
> This, I find completely unbelievable.

You both have anecdotes. Anecdotes don’t “cancel” each other out.

Here’s a third one. In some of the code reviews I’ve encountered that AI gives a lot of feedback, it’s just providing noise. Things that should be ignored or when following the feedback causes more harm which requires more token to “fix” later on. That can also happen. Sometimes the thing it spits out goes against the common sense, and sometimes it works very well.

> So, to anyone who insists on trying to keep up with AI, I say: good luck.

This I agree with, for a different reason. It’s like trying to swim in a sea of honey and trash mix. It’s exhausting.

[−]abuani · 2026-10-10 Sat 22:14 UTC · link
> Here’s a third one. In some of the code reviews I’ve encountered that AI gives a lot of feedback, it’s just providing noise

Something I've found fun is seeing how long it takes for an llm review tool to come back satisfied with a PR. Think 100 lines of code changed, nothing terribly significant, but also not trivial. I'll have a local Claude session setup to babysit the PR and wait for feedback, accept all the recommendations, push the change up and request a review. I cap the number of iterations at 10 just so I'm not blowing a stupid amount of money. I've yet to come up with a PR where the llm reviewer is satisfied with the changes and has _no feedback_.

So where's the reasonable cutoff point for llm based reviews?

[−]bgoated01 · 2026-10-11 Sun 00:20 UTC · link
Huh, I've had many times when `codex /review` comes back satisfied on the first shot, both with handwritten and LLM-assisted PRs.

We use only one round of LLM review, and have discussed as a team still assessing recommendations, not just accepting everything blindly. So somewhere in the (0,1] rounds of review. Sounds like our preferred ratio of human to LLM involvement is different than yours, though.

[−]enraged_camel · 2026-10-10 Sat 22:32 UTC · link
>> Here’s a third one. In some of the code reviews I’ve encountered that AI gives a lot of feedback, it’s just providing noise.

This is completely normal if you haven't written your own custom skill with instructions on what types of issues the AI should emphasize, which ones would be considered nits and which ones aren't a problem at all.

In our repos we use a classification system: blocker, should-fix and nit. Each one has specific definitions, criteria and examples encoded in the skill file. When the time comes to review a PR, agents invoke the skill, and frankly do a stellar job. A human then reads each finding, asks the AI follow-up questions and makes the final decision in terms of whether the finding goes in the PR review.

The reason I know this works is that we have one guy on the team who does not use this skill, and blindly throws his agents at PRs. And the results are exactly as you describe.

[−]LoganDark · 2026-10-10 Sat 23:49 UTC · link
> This is completely normal if you haven't written your own custom skill with instructions on what types of issues the AI should emphasize, which ones would be considered nits and which ones aren't a problem at all.

"You're holding it wrong" :)

[−]lmz · 2026-10-11 Sun 04:42 UTC · link
> "You're holding it wrong" :)

It's a general purpose tool. There are different ways of applying it, with different results.

[−]geraneum · 2026-10-11 Sun 05:05 UTC · link
> This is completely normal if you haven't written your own custom skill with instructions

This line of thinking comes from the assumption that LLM are somehow infallible and can’t be wrong, which is obviously not the case.

Writing skills and other magic incantations is the first thing that comes to mind of anyone who sees the review results for the first few times. What makes you think we didn’t do that?