btay.io/wiki

Estimating an LLM bill

The token arithmetic every platform bills on — where the money actually goes, and which levers move it.

Updated 18 d ago

Every managed AI platform bills the same underlying thing in slightly different currency: tokens in, tokens out, plus whatever infrastructure you left running. The published rates change constantly and the structure does not, so the useful thing to carry around is the structure. This page is the arithmetic behind the platform pages, and it should still hold when every number on AWS Bedrock and Snowflake Cortex has been revised twice.

The unit

A token is a subword chunk, not a word. The working conversion is 4 characters per token, or roughly 750 English words per 1,000 tokens. A paragraph is 100 to 150. A page is 500 to 700.

You are metered on two streams with different prices:

StreamIsTypically
Inputyour prompt, system instructions, retrieved context, the row you passed incheaper
Outputwhat the model generates3x to 8x the input rate

That asymmetry is the single most useful fact here, because it means the shape of your workload decides your bill more than its volume does.

Workload shape beats volume

Two jobs, same token count, very different cost:

JobInputOutputDominated by
Summarize a document5,000200input, so cheap per call
Draft a document2005,000output, so 3x to 8x more
Classify a row3005input, and trivially cheap
Agent loop, 10 turnsgrows every turnsmall each turnaccumulated input

The agent row is the one that surprises people. Each turn resends the whole conversation, so input grows roughly quadratically with turn count even though each individual reply is short.

The four levers

In descending order of how much they move the bill.

1. Model choice. The spread between the cheapest usable model and the most capable one is one to two orders of magnitude. Most classification, extraction, tagging and summarization work does not need a frontier model, and routing by task complexity is usually the largest saving available. Measure quality on your own data before assuming you need the expensive one.

2. Output length. Since output costs several times input, constraining the response is a direct cost control. Ask for JSON with a fixed schema, cap the list length, forbid the preamble. "Answer in one sentence" is a budget decision.

3. Prompt caching. A stable prefix, a system prompt or a shared instruction block, can be cached and re-read at roughly a tenth of the input rate. There is usually a write premium on first use, so it pays off after about two hits. For anything agentic with fixed instructions this stops being an optimization and becomes the architecture.

4. Retrieval discipline. Every retrieved chunk you stuff into context is input you pay for on every call. Retrieving 20 chunks when 5 would do is a 4x input bill for the same answer, and it often makes the answer worse.

What is not billed per token

This is where the surprise charges live, because the meter runs whether or not anyone is using the thing.

Charged onExamplesRuns when idle
Stored index sizevector and search indexes, billed per GB per monthyes
Provisioned throughputreserved model capacityyes
Per document or pageOCR and layout parsingno
Per messagesome managed query interfacesno
Compute hoursfine-tuning, batch jobs, notebooksonly while running
Egress and storagethe documents behind the indexyes

An index you built for a pilot that nobody uses is the classic recurring line item. It has no token cost at all, which is exactly why it never shows up in a token-based estimate.

Estimating before you build

Work in this order:

  1. Count the calls. Rows, documents, users, turns per session. This is a product question, not a modelling one, and it is where most estimates are wrong.
  2. Measure one real call. Take an actual prompt with actual retrieved context, run it, and read the token counts back from the API. Do not estimate from character counts of a prompt you have not filled in.
  3. Multiply, then apply the asymmetry. Input and output separately, at their own rates.
  4. Add the idle infrastructure. Index size, provisioned capacity, storage.
  5. Then apply the levers. Caching hit rate, model routing, output caps.

Steps 1 and 2 dominate the accuracy of the answer. Steps 3 to 5 are arithmetic.

Interactive against batch

The same work costs differently depending on when it runs.

Interactive workloads pay for latency: you want a fast model, you cannot batch, and you often over-retrieve to avoid a follow-up round trip. Batch workloads can use a slower and cheaper model, deduplicate aggressively, and cache heavily because they see the same prefix thousands of times in a row.

If a workload does not need to answer a human in real time, saying so is a cost decision.

Watch the meter from the first day

Every platform exposes per-call usage somewhere, and the useful dimensions are always the same: model, function, user, and the job that made the call. Three controls are worth wiring up before scale rather than after a bill:

  • A daily aggregate with a threshold that alerts
  • Per-user or per-team attribution, so a runaway job has an owner
  • A cap on any single job, so a bad loop terminates rather than completes

The failure mode this prevents is discovering a 59x model-choice mistake at the end of a billing period rather than in the first hour.

Gotchas

  • A per-row cost of $0.0002 is a real number when there are ten million rows. Always carry the estimate through to the full batch before deciding something is cheap.
  • Retries are billed. A pipeline that retries three times on failure has three times the cost ceiling, and a bad prompt fails repeatedly.
  • Long-context variants are usually a separate, dearer rate. Selecting one you do not need can silently double a workload.
  • Fine-tuning has two prices, one to train and a higher one to infer against the tuned model. The training cost is the one people remember and the smaller of the two at scale.
  • An index bills while idle, so a decommissioned project needs decommissioning. Deleting the application does not delete the index.

See also

Related pages

On this page