The best models have O(1k tokens per second). A billion tokens, sequentially, takes 1-2 weeks. 500B takes 500x that ... not fast. The problem might not have required 500B tokens, but if it needed anything within a couple orders of magnitude then something like the given approach (or anything else yielding equivalent results in exchange for parallelism) was mandatory.
That's plausibly true. I've definitely seen absurd token costs abused and wasted. I've also seen simple problems require absurd token counts regardless of prompt quality. Which factors made this problem require 1000x fewer tokens than they used?
When you see these absurd numbers, they are re-counting cached input for every turn. So a simple tool call when your context size is at 500k counts as another 500k to the sum.
Total Opus 5.5 token usage on OpenRouter last week is 5000B, presumably not counting cached input.
Total Opus 5.5 token usage on OpenRouter last week is 5000B, presumably not counting cached input.