Skip to content
AIAn Alian Software company
AI10 min read

Why Your AI Bill Tripled After Launch: The Cost Levers Nobody Prices Into the Proposal

The proof of concept cost pennies a conversation. Ninety days after launch, the invoice is three times the forecast — at the same traffic. Nothing broke; the system simply started behaving like production: longer prompts, fatter context, agents that take extra steps, and a flagship model answering questions a small one could handle. Here are the six leaks behind almost every AI cost overrun, the six levers that close them, and the pre-launch review that keeps the first invoice from starting an argument.

  • agents
  • engineering
  • roi

Every AI proposal prices the build with care — discovery, integration, evals, the launch week — and then treats running cost as a footnote: a token estimate multiplied by forecast volume, usually based on whatever the proof of concept consumed. Then the system goes live and the footnote becomes the biggest line on the invoice. The CFO asks why the assistant costs three times what was quoted, traffic is roughly where everyone said it would be, and nobody can point to the change that did it. That's because there wasn't one change. There were six small ones, each invisible in a demo and each billed by the token on every call. This post is the cost discipline for production AI: where the money leaks, which levers close each leak, and how to price the run — not just the build — before anyone signs.

The core argument in one paragraph: an LLM system is billed for every token it reads and writes, and production makes it read far more than any demo did — a system prompt that grew from 600 to 3,000 tokens, retrieval that pulls eight chunks "to be safe," conversation history re-sent on every turn, agents that make four model calls where the estimate assumed one. Output grows too, and output tokens typically cost several times more than input. None of this throws an error; all of it reaches the invoice. The fix is not a cheaper vendor. It's six engineering levers — caching, routing, batching, context trimming, agent budgets, and output limits — plus per-feature cost tracking, and every one of them can be scoped and priced before launch.

Why production costs more than the proof of concept (mechanically, not mysteriously)

Six behaviours drive almost every overrun we see, and the proof of concept exercises none of them:

System prompts only ever grow. Every edge case found in UAT becomes another paragraph of instructions — tone rules, refund policy, escalation logic, compliance wording — and every tool the agent can call adds its schema. A 600-token prompt becomes 3,000 tokens within weeks, and it is re-sent, in full, on every single call.

Retrieval is tuned for quality, never for cost. RAG pipelines get adjusted until answers are right, which usually means pulling more chunks, not better ones. Eight chunks of 300 tokens on every request is 2,400 tokens of context the model reads whether the question needed it or not.

Chat history compounds. Most chat products re-send the whole conversation each turn. A ten-turn support chat pays for turn one ten times — and long conversations are exactly the ones where users need help most.

Agents take steps the estimate never counted. An agent that plans, calls a tool, checks the result, and calls another tool makes four model calls where the proposal assumed one. Add retries on malformed JSON, timeouts, and failed validations, and the call count per task becomes a distribution with a long, expensive tail.

One flagship model does every job. Classifying an intent, extracting an order number, and writing a careful reply to an angry customer are priced identically when they all hit the most capable model. Most production traffic is the first two kinds of work.

Output runs long. Output tokens typically cost several times more than input tokens, and models default to thorough. A model that writes 400 words where 120 would do is the most expensive habit in the stack — and often the least helpful one too.

What a 3.3× overrun looks like, line by line

Take a representative customer-support agent — an illustrative composite, not a single client — and assume output tokens cost about five times input tokens, a common ratio across major providers. The proposal estimated three model calls per conversation, each carrying a 600-token system prompt, 800 tokens of retrieved context, and the customer's message: roughly 6,000 input tokens and 900 output tokens per conversation. In input-token terms, that's about 10,500 cost units per conversation.

Month two looks different. The system prompt is 3,000 tokens after UAT. Retrieval returns eight chunks. One extra tool step means four calls instead of three. Input per conversation is now around 25,600 tokens and output has doubled to 1,800 — roughly 34,600 cost units, or 3.3× the quote at identical traffic. Note where the growth came from: input grew more than fourfold while output only doubled. Most of the overrun is context the model re-reads, not text it writes — which is good news, because re-read context is the cheapest thing to fix.

Apply four of the levers below — cache the static prompt, rerank retrieval down to four chunks, summarise history and cap output length — and the same conversation drops to about 15,400 cost units, or 1.5× the original estimate, without changing the model or removing a feature. Routing simple turns to a smaller model closes most of what remains.

Levers 1–3: cache what repeats, route by difficulty, batch what can wait

Lever 1: Prompt caching — the closest thing to free money in this post. The system prompt, tool definitions, and policy text are identical on every call, so stop paying full price to re-read them. Major providers bill cached input at a steep discount — up to around 90% off on some models — once the same prefix is reused within the cache window. The engineering is mostly ordering: static content first, variable content (retrieved chunks, history, the user's message) last. Put a timestamp or user ID at the top of the prompt and the cache never hits; that one detail is the most common reason teams "enable caching" and see no change on the invoice.

Lever 2: Model routing — the right model for each job, not the best model for every job. Split traffic by difficulty. Intent classification, field extraction, language detection, and short FAQ answers rarely need a flagship model; a small model handles them at a fraction of the price and usually faster. Reserve the large model for reasoning, judgement calls, and replies where tone and accuracy carry real risk. A router can be a few rules, a cheap classifier, or a small model deciding when to escalate — but every route change goes through the eval set first, because a cheaper model that quietly gets 4% of refund questions wrong isn't cheaper.

Lever 3: Batch processing — half price for patience. Anything that doesn't need an answer in seconds — overnight product descriptions, ticket tagging, document summaries, weekly reports, re-embedding a knowledge base — belongs on a provider's batch API, which most providers price at roughly half the standard rate. The trick is noticing which jobs are actually asynchronous. In most systems we audit, a surprising share of "real-time" calls are background jobs that were simply built on the synchronous endpoint because it was there first.

Levers 4–6: trim the context, budget every agent, cap the output

Lever 4: Context trimming — fewer, better tokens. Retrieval should return the chunks that answer the question, not every chunk that might. A reranker in front of the prompt lets you fetch wide and send narrow: pull twenty candidates, keep the best three or four. Summarise conversation history after a few turns instead of re-sending it verbatim. Strip HTML, navigation text, legal boilerplate, and duplicate passages before anything reaches the model. In the example above, this lever alone removed more tokens than any other change — and answer quality typically holds or improves, because the model has less noise to reason through.

Lever 5: Step and retry budgets — a ceiling on every loop. Every agent gets a hard limit on tool calls per task and retries per call, and every time it hits the limit, that's logged as an event worth reviewing. Without a budget, the cost of a task is unbounded: one malformed tool response, one ambiguous instruction, and an agent can loop until a timeout ends it. With a budget, a runaway task costs you one capped task, escalates to a human, and shows up in Monday's review instead of next month's invoice. Fix validation errors at the source, too — a schema the model keeps getting wrong is a retry tax on every call.

Lever 6: Output limits — shorter is cheaper, and usually better. Set a max-token ceiling per response type, and state the length you want in the prompt itself, because the ceiling only truncates while the instruction actually shapes the answer. Where a parser rather than a person reads the output, ask for structured JSON with no preamble. Because output is the most expensive token class, this lever moves the invoice more than its simplicity suggests — and customers rarely complain that a support answer got to the point faster.

Cost observability: you can't cut what you can't attribute

A provider dashboard tells you that spend went up. It never tells you why. The minimum viable setup is to tag every model call with four things — the feature that made it, the customer or tenant it served, the model and prompt version used, and its input, cached, and output token counts — and send them to the same place you already keep traces. With those fields you can answer the questions that actually drive decisions: which feature costs the most per use, which customer is unprofitable, which prompt change doubled spend last Tuesday, and what one resolved support ticket really costs compared to a human handling it.

Then alert on unit cost, not total spend. Rising cost per conversation is an early warning that a prompt grew, retrieval widened, or an agent started looping. Rising total spend with flat unit cost is just growth — the good kind of invoice. This is the same discipline from our staging post applied to money: every model or prompt change gets a cost delta alongside its quality delta, and the cost of yesterday's surprise becomes a regression check for tomorrow's release.

The pragmatic version (a one-hour cost review before launch)

For an SMB running a support agent and a few automations, the whole discipline compresses to a single meeting before go-live, with the buyer and the build team in the same room. Estimate cost from production-sized prompts — tool schemas and policies included — not the demo prompt. Agree on how many model calls a typical task makes, and what the worst case is. Confirm the static part of every prompt is ordered for caching. List which request types can run on a smaller model, and check the eval set says quality holds. Move every job that can wait to batch. Write down how many chunks retrieval returns and how history is trimmed. Give every agent a step limit and a retry limit. Set output ceilings per response type. And make sure every call is logged by feature, customer, and model, with an alert on unit cost.

Any question in that list without an answer is the line item most likely to surprise you. Answer them before launch and the cost estimate stops being a guess based on a demo; it becomes a number with the same assumptions written beside it that the invoice will later test.

The strategic point underneath: AI is the first category of business software where the run cost scales directly with how much people use it. A busy assistant can cost more to operate in its first year than it cost to build — which means pricing the build carefully and the run casually is pricing the smaller number carefully. A tripled invoice isn't evidence the project failed; it's evidence the system is finally doing real work under real conditions the demo never simulated. That's fixable, and the fixes are boring in the best way: cache what repeats, route what's simple, batch what can wait, trim what isn't needed, cap what can loop, and measure everything per feature.

Every system we ship includes this cost layer — caching, routing, agent budgets, and per-feature cost tracking — because the first invoice should confirm the forecast, not start an argument.

Monthly briefing

One short email a month — what we shipped, what we learned, the patterns we'd recommend (and skip). No fluff.

Got a problem like this?

Describe it in the hero — our agent will scope a solution and tell you what a real build would look like.