The Ledger Gap

We track exactly what every AI call costs us. One month, our numbers said one thing. Google's real bill said something a lot bigger. This is the story of where the missing money was hiding — and it turned out to be something the AI was doing that nobody had ever told us about.
Every time our app calls an AI model — writing a lesson plan, answering a chat question, building a Teach Mode slide — we log exactly what it costs. That's how we can say, confidently, "one lesson plan costs about ₹2.80." Cheap, and precise. So when the real bill from Google showed up nearly double what we'd logged for the month, that wasn't supposed to be possible. We track everything. Except, it turned out, we didn't.
If what you tracked and what you got billed don't match, the bill is right. Your tracking has a hole in it somewhere.
Before we get to the big one: we did find one small leak early on — a couple of background jobs (like our image search) that skipped our tracking system entirely, so their tiny bit of cost was quietly invisible. Worth patching, but it wasn't the real story. It barely moved the number. The real money was hiding somewhere much sneakier.
THE CLUE
A screenshot where the numbers didn't add up
Then one afternoon, we observed something — one single AI call, the kind our app makes dozens of times a day without anyone giving it a second look. This one just felt a little off, in a way that's hard to put into words until you actually look at the numbers underneath it.
Add the input and the output together — 3.6K and 2.6K — and you get 6.2K. The total on screen said 9.6K. Somewhere, 3.4K tokens, more than a third of the whole call, had slipped in without ever showing up on either line above it. And right underneath that, the price tag was doing something strange: $0.0011 plus $0.0065 landed on $0.0076 exactly. Not close — exact. Whatever was quietly using those extra tokens wasn't costing anything on paper. Which almost certainly meant it was costing something in real life, and paper just hadn't noticed yet.
THE CAUSE
Gemini was thinking. We just weren't listening.
That's when it clicked: Gemini 2.5 doesn't just answer straight away. Before it writes the reply you actually see, it works the problem out first, in its own words — a kind of scratchpad — and then throws that scratchpad away before showing you anything. We'd never really had a reason to think about this before. Now we did.
Google still counts that scratchpad. It's baked into the number called "total tokens." It is not baked into the number called "output tokens" — and output tokens was exactly the field our cost math had been reading as the output, this whole time. Every bit of thinking Gemini did was real, billed, and completely invisible to us.
Which begged the obvious question: why was there so much of it? Poking around the code that talks to Gemini turned up a line that was supposed to answer exactly that — a cap on how much thinking Gemini was even allowed to do:
Except the condition guarding it could never actually be true. AI_PROVIDER.apiKey is Gemini's own key, and it's always sitting there the moment Gemini is the active model — so !apiKey is always false, and this line had never once fired since the day it was written. There had never been a cap. Gemini had been thinking as much as it wanted, on every call, for as long as this feature had existed.
To be sure that really was it, we asked Gemini the exact same tiny question twice, live — once with thinking left on, once switched off — and read back exactly what Google's own API said:
With thinking left on, the answer came back correct — and the total quietly carried 190 tokens that neither the prompt nor the reply could account for. Turn thinking off, and prompt plus reply landed on the total exactly, down to the token. There it was, sitting in two rows of a table: the same gap from that screenshot, now happening on command.
FIX
We stopped trusting "output tokens" on its own. Now, whenever the total comes back bigger than prompt + reply, we treat that extra chunk as output too — because that's genuinely what got billed, thinking included.
And in the end, it was a subtle difference in definition that nearly cost us double, reminding us that with AI, sometimes you have to read between the tokens.



