Hacker News

Favorites Setup
Comment by RugnirViking | original | Grieving the loss of details
[−]RugnirViking · 2026-10-10 Sat 22:05 UTC · link
> You could potentially ask an LLM to explain it

You can ask frontier models with all the bells, whistles, language servers etc to give you every example of X. And it gives you 12 examples and says that's all of them. When you know damn well there are at least 40 (but not exactly how many). So you say no, you know it has missed some, such as X23, and X27. So it goes away and comes back again and says yes, there are 41 X, here they are. How much digging should you do to see if its right?

I have many times gone looking myself to find that it has still missed some. Asked it to go check for those, to look harder it goes oopsie and says now im sure ive gotten all of them (has it?)

All this to say, I have been burned, repeatedly, multiple times a day, for the last year+. I still use these tools, but I truly cannot understand how some people treat them as oracles that know everything about our codebases

[−]cjkaminski · 2026-10-11 Sun 00:33 UTC · link
Excellent points. I agree that frontier models get things wrong all the time. They are NOT oracles. I believe that we need to encourage our peers to become sophisticated operators of these systems, instead of treating the output like Moses and the 10 Commandments. Perhaps we need more specialists, because the major AI labs lack the incentives to venture down the long tail of knowledge. Agents can be a piece of the learning puzzle, despite their imperfections. Again, thanks for your reply. It gave me lots to think about.
[−]manmal · 2026-10-11 Sun 09:59 UTC · link
You need to define the goal in a way that you will trust the completion. If you are not sure how you yourself would verify that there are 41, then you are in trouble. Verification must be deterministic, or it's worthless.

What the large models do really well nowadays, is that they won't lie to you if your deterministic verification fails. If you tell an OpenAI or Anthropic model that they need to run a certain `grep` or search or whatever command to verify, then they will do it. I haven’t seen them lie about this for a year, and trust them in this.

[−]RugnirViking · 2026-10-11 Sun 11:01 UTC · link
Surely you can see that for basically every example of this sort of problem, fully defining a deterministic check is the same as finding them all?

Like you're telling me if I had a script that printed all X, and it's my responsibility to ensure it has no bugs, then the agent could tell me all X and I could trust it?

This is not helpful at all? I am capable of running scripts myself and using the output directly.

[−]manmal · 2026-10-11 Sun 11:21 UTC · link
> fully defining a deterministic check is the same as finding them all

That’s like saying mathematics is worthless and we have to resort to finger counting?

[−]RugnirViking · 2026-10-11 Sun 12:19 UTC · link
No? It's like saying "why bother doing maths, a calculator can figure out the solution to any question you have if you just put the right equation in" there might be more to maths, it turns out. Most of it lying inside of finding "the right equation"