You need to define the goal in a way that you will trust the completion. If you are not sure how you yourself would verify that there are 41, then you are in trouble. Verification must be deterministic, or it's worthless.
What the large models do really well nowadays, is that they won't lie to you if your deterministic verification fails. If you tell an OpenAI or Anthropic model that they need to run a certain `grep` or search or whatever command to verify, then they will do it. I haven’t seen them lie about this for a year, and trust them in this.
Surely you can see that for basically every example of this sort of problem, fully defining a deterministic check is the same as finding them all?
Like you're telling me if I had a script that printed all X, and it's my responsibility to ensure it has no bugs, then the agent could tell me all X and I could trust it?
This is not helpful at all? I am capable of running scripts myself and using the output directly.
No? It's like saying "why bother doing maths, a calculator can figure out the solution to any question you have if you just put the right equation in" there might be more to maths, it turns out. Most of it lying inside of finding "the right equation"
What the large models do really well nowadays, is that they won't lie to you if your deterministic verification fails. If you tell an OpenAI or Anthropic model that they need to run a certain `grep` or search or whatever command to verify, then they will do it. I haven’t seen them lie about this for a year, and trust them in this.
Like you're telling me if I had a script that printed all X, and it's my responsibility to ensure it has no bugs, then the agent could tell me all X and I could trust it?
This is not helpful at all? I am capable of running scripts myself and using the output directly.
That’s like saying mathematics is worthless and we have to resort to finger counting?