Token spend, rework, switching, hidden thinking and context re-read. Together they are the whole cost of finishing a piece of AI work.
Five Costs, Priced Before the Request Is Sent
Each cost counts money the others do not. Together they are the whole cost of finishing the work.
Token spend. What the provider bills, counted across every attempt a model will really need, not the sticker price of one request.
Rework. The time a person spends noticing a wrong answer and putting it right, at their real hourly rate. Usually the largest cost on short work, and never on the AI invoice.
Switching. Moving half-finished work to another model throws away the discount the provider gives for a conversation it already holds.
Hidden thinking. Some models reason privately and bill for thinking you never see. The cheapest price list can mean the most expensive run.
Context re-read. The whole conversation is charged again on every turn. On long sessions this is the biggest line on the bill.
Line by Line
Every saving comes apart line by line on the receipt, so a buyer can follow any figure to its source. When routing up spends more on tokens, that line is printed negative. See CompletionPrism™.