Hacker News

Favorites Setup
Comment by isubkhankulov | original | 500B Tokens Later: Letting AI Agents Decompile a First-Person Shooter
[−]isubkhankulov · 2026-10-11 Sun 07:53 UTC · link
to clarify, I mean that the big model companies are charging consumers much less than equivalent API pricing but they're not actually losing money so its more of a steep at-cost discount for inference. It likely does not fully cover amortized R&D just like the other reply stated but it does still cover marginal inference cost (GPU/power/etc)

Uber and Lyft were paying drivers $X but charging users way less than $X so they were literally burning investor money to get market share.

[−]xienze · 2026-10-11 Sun 09:12 UTC · link
> It likely does not fully cover amortized R&D just like the other reply stated but it does still cover marginal inference cost (GPU/power/etc)

Why does this point come up over and over again, pretending that you can truly separate training and inference costs. Yes, they are separate things but the value OpenAI and Anthropic are presenting to the world is "we're the absolute best, no one else comes close." Well, to keep that up you can't just not train for extended periods of time. You have to keep that engine going non-stop when there's free Chinese models nipping at your heels. You can be profitable on "just inference" all you want but if training expenses dwarf that, you're not going to be profitable overall, and that's the bottom line.

[−]edg5000 · 2026-10-11 Sun 09:30 UTC · link
> does still cover marginal inference cost Simply comparing to the larger models on OpenRouter implies that the pure hosting costs (equipment + power + minimal overhead) still exceed plan pricing if we assume all users always use their full weekly allowance.

So my conclusion is that it only works because the majority of users doen't fully utilize their allowance. Last month I used almost nothing of my Claude 20x plan (did use Codex though).