The first agent bill is usually a surprise. A chat costs cents, so teams expect an agent to cost a few times more. Then a single task turns out to cost more than a dollar, and a thousand tasks a day is a five-figure monthly line item.
The reason is structural, and it means most of the bill can be removed without touching quality. This guide explains where an agent's tokens actually go, then walks through the levers that matter most in 2026 (prompt caching, model routing, batching and output control) using Anthropic's published prices and a cost model in Python you can run against your own numbers. In our modeled workload, caching alone cut the bill by 77%, and one common mistake made it 23% worse than no caching at all.
Why agent bills explode
A language model API is stateless. On every step of an agent loop, you send the entire conversation again: the system prompt, every tool definition, every earlier tool result and every earlier reply. A 15-step task doesn't send the prompt once. It sends it 15 times, and each copy is longer than the one before.
Take a realistic shape: a 12,000-token prefix (system prompt plus tool definitions), tool results of 1,500 tokens per step, and replies of 400 tokens. Over 15 steps, the model writes 6,000 output tokens but reads 402,000 input tokens. On Claude Opus 5.5 that is $1.61 of input against $0.12 of output, so 93% of the bill is re-reading the same context. (Multi-agent systems multiply this further; Anthropic measured about 15× the tokens of a chat. See one agent or many?)
That number tells you where to look. The most effective savings don't come from shorter answers. They come from not paying full price to re-read a prefix the model has already seen.
Lever 1: prompt caching
Prompt caching stores the processed prefix of your prompt so later requests that start with the same content read it from cache instead of processing it again. Anthropic's documentation (checked October 5, 2026) gives the pricing as multipliers of the normal input price:
- Cache write: 1.25× for the default 5-minute cache, or 2× for a 1-hour cache.
- Cache read: 0.1× on most models, and lower on the newest ones. Claude Opus 5.5 reads cached tokens at $0.20 per million, against $4 for normal input: 5% of the price.
Turning it on is now one field. With automatic caching, you add a single top-level cache_control to the request and the cache point moves forward as the conversation grows, which is exactly the shape of an agent loop:
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=1024,
cache_control={"type": "ephemeral"}, # automatic caching, 5-minute TTL
system=SYSTEM_PROMPT,
tools=TOOLS,
messages=conversation,
)
Four rules decide whether you actually get cache hits:
- Matching is an exact prefix. Cache hits require identical content up to the cache point, in the order
tools→system→messages. A change at any level invalidates that level and everything after it. - Put anything that changes at the end. The documentation's own "common mistake" example is a timestamp in a block that changes every request: you pay for a fresh cache write every time and never get a read.
- Keep tool definitions and settings stable. Changing tool definitions invalidates the whole cache; switching
tool_choice, toggling images, or on some models changing thinking or effort settings invalidates part of it. The docs also warn that some languages (Swift and Go, for example) randomize JSON key order, which silently breaks caching. - Short prompts don't cache. There is a minimum length per model, from 512 tokens on the newest models to 4,096 on some older ones. Agent prefixes are almost always above it.
Caches are isolated per workspace and never shared between organizations, and cached prefixes also improve time-to-first-token, so caching makes agents faster as well as cheaper.
See the numbers: a cost model you can run
The script below prices one agent task under several setups using the published per-million-token prices. The bill() function takes the same fields the API returns in response.usage (input_tokens, cache_creation_input_tokens, cache_read_input_tokens, output_tokens), so you can also point it at real responses from your logs. Create agent_cost.py:
"""What an agent run really costs, and what caching, routing and batching save.
Prices are USD per million tokens from Anthropic's pricing page (checked
2026-10-05). The workload is an assumption: change it to match your own logs.
`bill()` takes the same usage fields the API returns, so you can point it at
real responses too.
"""
PRICES = { # input, 5-min cache write, cache read, output
"claude-opus-5-5": {"in": 4.00, "write": 5.00, "read": 0.20, "out": 20.00},
"claude-sonnet-5-5": {"in": 2.00, "write": 2.50, "read": 0.20, "out": 10.00},
"claude-haiku-4-5": {"in": 1.00, "write": 1.25, "read": 0.10, "out": 5.00},
}
PREFIX = 12_000 # system prompt + tool definitions, identical on every call
TOOL_RESULT = 1_500 # tokens a tool returns per step
REPLY = 400 # tokens the model writes per step (tool call or answer)
STEPS = 15 # model calls per task
def bill(usage: dict, model: str) -> float:
"""Dollar cost of one API response, from its usage block."""
p = PRICES[model]
return (usage.get("input_tokens", 0) * p["in"]
+ usage.get("cache_creation_input_tokens", 0) * p["write"]
+ usage.get("cache_read_input_tokens", 0) * p["read"]
+ usage.get("output_tokens", 0) * p["out"]) / 1_000_000
def run_task(model: str, caching: bool, stable_prefix: bool = True) -> float:
"""Simulate one agent task: each step re-sends the whole conversation so far."""
total, history, cached = 0.0, 0, 0
for _ in range(STEPS):
history += TOOL_RESULT
prompt = PREFIX + history
if not caching:
usage = {"input_tokens": prompt}
elif not stable_prefix: # e.g. a timestamp at the top of the system prompt
usage = {"cache_creation_input_tokens": prompt} # rewrites everything, never reads
else: # automatic caching: read what was cached last call, write the new tail
usage = {"cache_read_input_tokens": cached,
"cache_creation_input_tokens": prompt - cached}
cached = prompt
usage["output_tokens"] = REPLY
total += bill(usage, model)
history += REPLY
return total
def main():
tasks = 1_000 # per day
opus = run_task("claude-opus-5-5", caching=False)
rows = [
("Opus 5.5, no caching", opus),
("Opus 5.5, cache busted by timestamp", run_task("claude-opus-5-5", True, stable_prefix=False)),
("Opus 5.5, prompt caching", run_task("claude-opus-5-5", True)),
]
# Routing: a cheap model classifies each task; 70% are routine and go to Sonnet.
router = bill({"input_tokens": 800, "output_tokens": 10}, "claude-haiku-4-5")
routed = router + 0.7 * run_task("claude-sonnet-5-5", True) + 0.3 * run_task("claude-opus-5-5", True)
rows.append(("Caching + routing (70% to Sonnet 5.5)", routed))
# Work that can wait (evals, backfills, nightly reports) can use the Batches API at 50% off.
# Batch cache hits are best-effort, so show the range: full hits vs none at all.
no_hits = router + 0.7 * run_task("claude-sonnet-5-5", False) + 0.3 * run_task("claude-opus-5-5", False)
rows.append(("Routing + batch, cache hits as above", routed * 0.5))
rows.append(("Routing + batch, zero cache hits", no_hits * 0.5))
print(f"{'setup':<40}{'$/task':>10}{'$/month':>12}{'vs baseline':>13}")
for name, cost in rows:
print(f"{name:<40}{cost:>10.3f}{cost * tasks * 30:>12,.0f}{cost / opus - 1:>+13.0%}")
if __name__ == "__main__":
main()
Running python agent_cost.py printed (monthly figures assume 1,000 tasks a day):
setup $/task $/month vs baseline
Opus 5.5, no caching 1.728 51,840 +0%
Opus 5.5, cache busted by timestamp 2.130 63,900 +23%
Opus 5.5, prompt caching 0.393 11,786 -77%
Caching + routing (70% to Sonnet 5.5) 0.282 8,447 -84%
Routing + batch, cache hits as above 0.141 4,223 -92%
Routing + batch, zero cache hits 0.562 16,861 -67%
What each row teaches:
- Caching is the big lever: −77%. After the first step, almost all input is a cache read at 5% of the price. You pay the 1.25× write premium only on the new tail of each step.
- A busted cache is worse than none: +23%. With a timestamp at the top of the system prompt, every step writes the whole context at 1.25× and never reads it back. This is the most common reason teams "turn on caching" and see their bill go up.
- Routing adds a smaller gain: −84% in total. It helps less than you might expect, because cache reads cost $0.20 per million on both Sonnet 5.5 and Opus 5.5. What routing saves is the cache writes and, above all, the output tokens.
- Batching halves what is left, if hits hold. Batch discounts stack with caching, but cache hits in a batch are best-effort, so the honest answer is a range: from −92% with full cache hits to −67% with none.
The model is arithmetic on published prices with an assumed workload, not a measurement of any bill. Its value is that you can replace the five workload constants with numbers from your own logs and see which lever matters for you. We also ship test_cost.py, which checks the baseline against a hand calculation.
Lever 2: route each task to the cheapest model that can do it
Most agent traffic is not hard. Answering an order-status question, filling a form from an email or tagging a ticket doesn't need the strongest model. A small classifier call (in our model, about 800 input tokens on Claude Haiku 4.5, a fraction of a cent) can decide which tasks are routine and send them to a cheaper model.
- Route per task, not per step. Keeping a whole task on one model keeps its conversation, and its cache, in one place. Switching models mid-loop gives up the cached prefix you already paid to write.
- Route on evidence. Decide the split by running both models on a sample of real tasks with your eval suite, not by intuition. (Our guide to testing agents with pass^k covers how.)
- Escalate, don't guess. If the cheaper model fails a check or reports low confidence, retry the task on the stronger model. A few escalations cost less than running everything on the top model.
Lever 3: batch anything that can wait
Anthropic's Message Batches API charges 50% of standard prices for both input and output. A batch can hold up to 100,000 requests (256 MB), most finish in under an hour, the maximum is 24 hours, and results stay available for 29 days. That makes it a natural fit for work with no user waiting: nightly reports, document backfills, classification of a whole archive and, importantly, your own eval runs.
Batching and caching discounts stack. Because batch requests run asynchronously and in parallel, though, the documentation describes cache hits in batches as best-effort, with typical hit rates between 30% and 98% depending on traffic. Budget with the conservative end of our range until your own logs show the real hit rate.
Lever 4: control output, the most expensive token
Output tokens cost five times as much as input on current Claude models ($20 against $4 per million on Opus 5.5). Two controls matter:
- The effort parameter. Setting
output_config.efforttolow,medium,high,xhighormaxtells supported models how many tokens to spend, across text, tool calls and thinking. Lower effort means fewer tokens, faster responses and lower cost, with reduced capability on hard problems. Since changing effort can invalidate the cache on some models, choose it per workload rather than per call. - Ask for less. Request structured, compact tool calls and final answers. A worker that returns a 200-token summary instead of a 2,000-token dump saves output now and input on every later step.
Lever 5: stop the context from growing
In the model above, each step adds 1,900 tokens that every later step re-reads. Trimming that growth compounds:
- Return less from tools. Filter search results, select only needed columns, and truncate logs before they enter the conversation.
- Move long-term facts out of the context. Store them in a memory tool and load only what the current step needs. (See giving your agent memory.)
- Load tools on demand. Every tool definition sits in the prefix of every call. If an agent has dozens of tools, expose only the ones the current task needs.
Measure it, or it will drift
Cost problems usually come back quietly: someone adds a date to the system prompt, reorders tools or changes a setting, and the cache hit rate falls without any error. Log the four usage fields for every call, compute dollars with a function like bill(), and track one ratio: cache-read tokens divided by total input tokens. In a healthy agent loop it should be high and stable. If it drops after a deploy, you have found your regression before the invoice does.
Bottom line
An agent's bill is dominated by re-reading its own context, so the order of work is clear. First, make the prefix stable and turn on caching; for most agents this is the largest single saving, and getting it wrong costs more than doing nothing. Then route routine tasks to cheaper models, send waiting work through batches, keep output and context lean, and watch the cache hit rate like any other production metric. Check prices on the official page before you plan a budget: they change, and the newest models have changed the cache-read math in your favor.
Sources: Anthropic Claude Platform documentation: Pricing, Prompt caching, Batch processing and Effort pages (checked 2026-10-05); Anthropic Engineering, "How we built our multi-agent research system" (June 13, 2025). Code tested on 2026-10-05 with Python 3.13. Costs are computed from published list prices with an assumed workload; they illustrate the levers and are not a measurement of any real bill. Prices change, so verify them before budgeting.