Estimating an LLM bill
The token arithmetic every platform bills on — where the money actually goes, and which levers move it.
Every managed AI platform bills the same underlying thing in slightly different currency: tokens in, tokens out, plus whatever infrastructure you left running. The published rates change constantly and the structure does not, so the useful thing to carry around is the structure. This page is the arithmetic behind the platform pages, and it should still hold when every number on AWS Bedrock and Snowflake Cortex has been revised twice.
The unit
A token is a subword chunk, not a word. The working conversion is 4 characters per token, or roughly 750 English words per 1,000 tokens. A paragraph is 100 to 150. A page is 500 to 700.
You are metered on two streams with different prices:
| Stream | Is | Typically |
|---|---|---|
| Input | your prompt, system instructions, retrieved context, the row you passed in | cheaper |
| Output | what the model generates | 3x to 8x the input rate |
That asymmetry is the single most useful fact here, because it means the shape of your workload decides your bill more than its volume does.
Workload shape beats volume
Two jobs, same token count, very different cost:
| Job | Input | Output | Dominated by |
|---|---|---|---|
| Summarize a document | 5,000 | 200 | input, so cheap per call |
| Draft a document | 200 | 5,000 | output, so 3x to 8x more |
| Classify a row | 300 | 5 | input, and trivially cheap |
| Agent loop, 10 turns | grows every turn | small each turn | accumulated input |
The agent row is the one that surprises people. Each turn resends the whole conversation, so input grows roughly quadratically with turn count even though each individual reply is short.
The four levers
In descending order of how much they move the bill.
1. Model choice. The spread between the cheapest usable model and the most capable one is one to two orders of magnitude. Most classification, extraction, tagging and summarization work does not need a frontier model, and routing by task complexity is usually the largest saving available. Measure quality on your own data before assuming you need the expensive one.
2. Output length. Since output costs several times input, constraining the response is a direct cost control. Ask for JSON with a fixed schema, cap the list length, forbid the preamble. "Answer in one sentence" is a budget decision.
3. Prompt caching. A stable prefix, a system prompt or a shared instruction block, can be cached and re-read at roughly a tenth of the input rate. There is usually a write premium on first use, so it pays off after about two hits. For anything agentic with fixed instructions this stops being an optimization and becomes the architecture.
4. Retrieval discipline. Every retrieved chunk you stuff into context is input you pay for on every call. Retrieving 20 chunks when 5 would do is a 4x input bill for the same answer, and it often makes the answer worse.
What is not billed per token
This is where the surprise charges live, because the meter runs whether or not anyone is using the thing.
| Charged on | Examples | Runs when idle |
|---|---|---|
| Stored index size | vector and search indexes, billed per GB per month | yes |
| Provisioned throughput | reserved model capacity | yes |
| Per document or page | OCR and layout parsing | no |
| Per message | some managed query interfaces | no |
| Compute hours | fine-tuning, batch jobs, notebooks | only while running |
| Egress and storage | the documents behind the index | yes |
An index you built for a pilot that nobody uses is the classic recurring line item. It has no token cost at all, which is exactly why it never shows up in a token-based estimate.
Estimating before you build
Work in this order:
- Count the calls. Rows, documents, users, turns per session. This is a product question, not a modelling one, and it is where most estimates are wrong.
- Measure one real call. Take an actual prompt with actual retrieved context, run it, and read the token counts back from the API. Do not estimate from character counts of a prompt you have not filled in.
- Multiply, then apply the asymmetry. Input and output separately, at their own rates.
- Add the idle infrastructure. Index size, provisioned capacity, storage.
- Then apply the levers. Caching hit rate, model routing, output caps.
Steps 1 and 2 dominate the accuracy of the answer. Steps 3 to 5 are arithmetic.
Interactive against batch
The same work costs differently depending on when it runs.
Interactive workloads pay for latency: you want a fast model, you cannot batch, and you often over-retrieve to avoid a follow-up round trip. Batch workloads can use a slower and cheaper model, deduplicate aggressively, and cache heavily because they see the same prefix thousands of times in a row.
If a workload does not need to answer a human in real time, saying so is a cost decision.
Watch the meter from the first day
Every platform exposes per-call usage somewhere, and the useful dimensions are always the same: model, function, user, and the job that made the call. Three controls are worth wiring up before scale rather than after a bill:
- A daily aggregate with a threshold that alerts
- Per-user or per-team attribution, so a runaway job has an owner
- A cap on any single job, so a bad loop terminates rather than completes
The failure mode this prevents is discovering a 59x model-choice mistake at the end of a billing period rather than in the first hour.
Gotchas
- A per-row cost of $0.0002 is a real number when there are ten million rows. Always carry the estimate through to the full batch before deciding something is cheap.
- Retries are billed. A pipeline that retries three times on failure has three times the cost ceiling, and a bad prompt fails repeatedly.
- Long-context variants are usually a separate, dearer rate. Selecting one you do not need can silently double a workload.
- Fine-tuning has two prices, one to train and a higher one to infer against the tuned model. The training cost is the one people remember and the smaller of the two at scale.
- An index bills while idle, so a decommissioned project needs decommissioning. Deleting the application does not delete the index.
See also
- AWS Bedrock AI Strategy for the platform-specific cost surface
- Snowflake Cortex: A Working Engineer's Guide for the token rates and monitoring queries
- What a security review asks for the other half of a platform decision
Related pages
- What a security review asksThe questions every enterprise AI platform gets asked, what a good answer looks like, and which ones are genuinely hard.
- AWS Bedrock AI StrategyConsolidated overview of Bedrock capabilities, use cases, architecture patterns, governance, and a sequenced project portfolio for a wealth management firm.
- Snowflake Cortex AI: A Working Engineer's GuideImplementation-level Cortex reference — product map, setup, AISQL functions, Cortex Search, the token cost model, monitoring queries, governance, and gotchas.
- Snowflake Cortex AI: The Big PictureA non-technical orientation to Cortex AI — what it is, why it matters strategically, where the real value is, what it costs, and what to be skeptical about.
- AWS Bedrock Use Cases for RIAsA catalog of 112 concrete Bedrock applications for a wealth management firm, grouped by function — document intelligence, data engineering, CRM, compliance, analytics, backfills, and agents.