Functional equivalence here, of course, depends on the completeness of the test suite, where byte-identical compiled artifacts does not.
(For example, your approach would not necessarily catch all the same overflow behaviors; the OP expressly claimed that "replicating all bugs" was also important, and many bugs are caused by certain overflow behaviors)
[−]SubiculumCode · 2026-10-11 Sun 02:37 UTC ·
link
Byte exact seems only of interest to preserve known bugs etc for cheats/shortcuts/etc.
That may be true, but I hate it when people repeat the false idea that functional equivalence requires only a test suite that has full branch/line coverage. Call me triggered :)
That said, I would probably follow this same approach if I were to do this, but with extensive randomized testing as well.
You can do the process in stages. Do the first decompilation mechanically (no LLM), use a SMT solver to show it builds to an equivalent binary to the original, and then use LLM to clean up the code into something idiomatic with the benefit of a correct binary built with the new toolchain. This helps when you want to port across languages or toolchains, and helps protect against toolchain bugs.
(For example, your approach would not necessarily catch all the same overflow behaviors; the OP expressly claimed that "replicating all bugs" was also important, and many bugs are caused by certain overflow behaviors)
That said, I would probably follow this same approach if I were to do this, but with extensive randomized testing as well.
So you're looking not just for functional but also dysfunctional equivalence %)