Hacker News

Favorites Setup
Comment by edg5000 | original | 500B Tokens Later: Letting AI Agents Decompile a First-Person Shooter
[−]edg5000 · 2026-10-11 Sun 03:41 UTC · link
I sense the approach overcomplicates things. I wonder how long this would have taken in a single session. Maybe this is actually a textbook example of something where subagents make sense, but when I first started LLMs I was often overcomplicating the workflow with all kinds of orchestration. Now I just use one agent, it better allows controlling the output even if the agent works slightly longer. Most time is spent by me writing prompts and reviewing work anyway (for me at least).
[−]hansvm · 2026-10-11 Sun 03:45 UTC · link
The best models have O(1k tokens per second). A billion tokens, sequentially, takes 1-2 weeks. 500B takes 500x that ... not fast. The problem might not have required 500B tokens, but if it needed anything within a couple orders of magnitude then something like the given approach (or anything else yielding equivalent results in exchange for parallelism) was mandatory.
[−]speedstyle · 2026-10-11 Sun 04:16 UTC · link
Yeah, I don't think the problem needed billions of tokens
[−]hansvm · 2026-10-11 Sun 04:54 UTC · link
That's plausibly true. I've definitely seen absurd token costs abused and wasted. I've also seen simple problems require absurd token counts regardless of prompt quality. Which factors made this problem require 1000x fewer tokens than they used?
[−]WASDx · 2026-10-11 Sun 08:13 UTC · link
When you see these absurd numbers, they are re-counting cached input for every turn. So a simple tool call when your context size is at 500k counts as another 500k to the sum.

Total Opus 5.5 token usage on OpenRouter last week is 5000B, presumably not counting cached input.