Hidden Thinking: The Reasoning Tokens on Your AI Bill You Never See

Reasoning models think privately before they answer, and bill for it. Why the model with the lowest list price can be the most expensive one to run.

Tokens You Pay For and Never Read

Many current AI models reason before they respond. They work through a problem in private, then produce the answer you see. That private reasoning is billed as output tokens, and on hard problems it can be several times longer than the visible answer. You pay for it whether or not you ever see it.

This is often a good trade. Reasoning improves accuracy on hard tasks, and fewer wrong answers means less rework. The problem is that it breaks the most common way companies compare models: the price list.

Why the List Price Misleads

Two models can carry similar prices per token and produce answers of similar length, yet cost very different amounts to run, because one of them thinks for ten times as long before answering. A team that chooses on list price may move work to a model that is cheaper per token and more expensive per answer.

When Reasoning Is Worth It

On a hard decision, a model that reasons and gets it right first time is usually the cheapest path to a finished result. On simple work, paying for long private reasoning adds cost and adds nothing. The answer depends on the step, which is why it cannot be settled by one company-wide rule.

Pricing It Before the Request

CompletionPrism™ counts hidden thinking as one of the five costs it prices before every request is sent, alongside token spend across every attempt, rework, switching and context re-reading. A model that looks cheapest on a price list does not win if its hidden reasoning makes it the most expensive way to finish the step. Each receipt labels the hidden thinking line as inferred, because it is worked out from the bill rather than observed directly.

What to Do Now

Compare models on cost per finished answer, not cost per token. Pull a week of usage, divide total output spend by the number of answers people actually kept, and see which models move.