Hacker News

Favorites Setup
Comment by aetherspawn | original | 500B Tokens Later: Letting AI Agents Decompile a First-Person Shooter
[−]aetherspawn · 2026-10-11 Sun 03:11 UTC · link
The reason this cost so much is because the AI has the ridiculous goal of getting identical assembly output.

The agents would have had to mess around with compiler versions, optimisation options, and the phase of the moon as well.

If you just went for functional equivalence, it would probably cost 10x or 100x less tokens.

Another false economy was using Sonnet instead of a more intelligent model like Sol 6.1 (1), which would have cost more per token, but is 100x or so better at reverse engineering and coding and therefore can chew through the source code much quicker and make fewer mistakes, meaning less work needing to be scrapped.

In my testing doing a similar task, I ran multiple sonnet for weeks and burnt through ~$1000 in tokens to get 20% completion and output that was pretty bad. After switching to Sol 6.1, it finished the whole task in around 2 days, cost around $50, and it did it with zero supervision and a single /goal.

(1): struggle to use Opus for reverse engineering, too many safeguards. OAI has virtually none, and uses way less tokens so is more economical.

[−]Gigachad · 2026-10-11 Sun 03:18 UTC · link
Identical output isn’t ridiculous. It’s pretty much a requirement to ensure the game actually is the same in every way. These decomp projects are pitched as a high performance alternative to emulation.

No one is going to use it if it’s a kind of close but not really reimplementation.

Testing functional equivalence is also pretty much impossible. How would you for example test the new one has exactly the same bugs which haven’t been discovered yet. Or doesn’t introduce new ones? This stuff matters for speed runners.

[−]aetherspawn · 2026-10-11 Sun 03:24 UTC · link
It’s ridiculous because something as simple as the compiler picking different registers is going to make zero functional difference but give a false negative on assembly compare.

Yet the C code can’t pick what registers to use, so the poor agent is probably shuffling the code around randomly for hours or days until it matches.

That’s probably why the agent dropped down into inline assembly in the first place (the author complained about this), because I bet it’s thinking trace was that this is futile.

Compilers themselves are not even deterministic and running them multiple times creates different assembly.

[−]boricj · 2026-10-11 Sun 08:20 UTC · link
> Yet the C code can’t pick what registers to use, so the poor agent is probably shuffling the code around randomly for hours or days until it matches.

Humans too do that for matching decompilation. Or at least I imagine so, given that I personally refuse to do that.

People have different goals and will use different techniques to achieve them. The video game reverse-engineering/decompilation community isn't a hive-mind, I went ahead and created ghidra-delinker-extension because I had my own ideas on how to do that.

[−]dezgeg · 2026-10-11 Sun 12:07 UTC · link
How do you know the original code didn't use some undefined variable on stack, but happened to consistently work because something else spilled a register containing "suitable" value?
[−]Neywiny · 2026-10-11 Sun 12:37 UTC · link
I think I disagree when it comes to bug reproduction. Doesn't seem I'm alone in that.
[−]stavros · 2026-10-11 Sun 10:13 UTC · link
Identical assembly output is the only way to get functional equivalence, otherwise you're getting equivalence in some percentage of situations, but not 100%. Your bug is my feature etc.