Short answer
Prompt caching cuts agent token costs by keeping the first N thousand tokens byte-identical across every loop iteration, so the provider's KV cache skips the expensive prefill step. LangChain measured 49–80% cost reductions across Anthropic, OpenAI, and Google models on real agent trajectories.
An agent that runs for 30 steps doesn't send one prompt. It sends 30, and each one is bigger than the last. The system prompt, the tool schemas, the retrieved context, and every prior tool call and observation ride along on every hop. By the time your agent finishes a non-trivial task, you've paid to process the same prefix dozens of times. Prompt caching for AI agents is the single biggest lever against that bill, and LangChain's public Deep Agents evaluation ran the suite across claude-haiku-4-5, gpt-5.4-mini, and gemini-3.5-flash on real trajectories to measure what the discount actually buys you end-to-end.
This piece is about where that saving comes from, how to design loops that capture it, and what to watch so the cache doesn't quietly stop hitting.
Why long agent loops burn tokens, and where caching helps
The 100:1 input-to-output ratio in agentic loops
Chat workloads and agent workloads have different cost shapes. In a chat, input and output are roughly comparable per turn. In an agent, they are not. Production agents routinely process roughly 100 input tokens for every output token, because every step re-ingests the entire trajectory to produce a short tool call or final answer.
So the bill is dominated by the input side. If you only compare per-token prices across models, you'll optimize the wrong thing; the real question is what fraction of those input tokens you're paying full freight on versus cache rate.
Why the tool schema and growing context get re-sent every step
LLM APIs are stateless. Each iteration of the ReAct loop has to include the full conversation history from step one, which for an agent means the system prompt, the tool definitions, any retrieved RAG context, and every prior (assistant → tool call → tool result) triple. Unlike a plain chat, the agent accumulates tool schemas and observations that quickly dwarf the user's original message.
The system prompt and the tool schemas are byte-identical across every step of a run. They are a perfect prefix for prompt caching, and if the cache hits, those tokens are served at a small fraction of base rate.
How prompt caching actually works
The KV cache: prefill, decode, and prefix matching
Inside the transformer, every input token's self-attention layers compute a pair of key and value tensors. That prefill is the expensive part of serving a long prompt. Prompt caching stores those KV tensors so a subsequent request with the same prefix can skip the computation entirely and jump straight to decoding the new tokens.
The prerequisite is universal: an exact prefix match. Change a single token in the cached portion and the match breaks. The model has to recompute from the first differing position onward. This is the mechanism behind cache invalidation, and it's the one thing that silently kills hit rate in production.
Prompt caching in provider APIs vs. raw KV cache
The research literature on KV cache reuse is broader than what providers expose. Academic systems like CacheBlend fuse cached chunks from arbitrary positions in the context rather than requiring a shared head, and recent benchmarks evaluate the full lifecycle of the KV cache under real multi-request load instead of single-request snapshots. The provider APIs are a constrained slice of that: strict prefix match, bounded TTL, per-provider billing. You don't get to splice cached chunks from the middle of two prompts together. You get to reuse a shared head.
That constraint is what makes prompt design matter so much. So prompt design is about making the head as long and stable as possible, and arranging anything that varies to sit after it.
Where the savings come from in long loops
The cost math: cache reads at a fraction of base, writes at a small premium
The economics across providers converge on roughly the same shape:
| Provider | Cache read | Cache write | Default behavior |
|---|---|---|---|
| OpenAI | discounted up to 95% | 1x | Automatic |
| Anthropic | ~0.1x base | ~1.25x | Explicit cache_control breakpoints |
| Google Gemini | Discounted (varies) | — | Implicit + explicit modes |
OpenAI's docs state cached input is discounted up to 95% depending on the model; check the current pricing table for yours. Anthropic charges a modest premium to write a cache entry, then reads at roughly 0.1x base for the lifetime of that entry. Independent analyses of Manus's architecture note cached tokens cost ~90% less than uncached ones in their stack, a figure consistent with Anthropic's own pricing examples.
The break-even is immediate. One re-read more than repays the write premium: at Anthropic's rates a write plus a single read costs about 1.35x base, against 2x for sending the same prefix twice uncached. In an agent loop that runs 20 steps against a 15k-token system-and-tools prefix, you pay the write premium once and the read discount 19 times.
What the Deep Agents eval actually measured
LangChain ran the Deep Agents evaluation suite across three mid-tier models (claude-haiku-4-5, gpt-5.4-mini, and gemini-3.5-flash) specifically to see what prompt caching saves on real agent trajectories rather than on a feature table. On real agent trajectories caching cut token cost by 49 to 80%: 80% for gpt-5.4-mini on automatic longest-prefix caching, 77% for claude-haiku-4-5 with explicit breakpoints, and 49% for gemini-3.5-flash on implicit caching. The spread is the point. Feature tables tell you what's possible; only running the eval across providers shows what lands on multi-step runs.
The gap between the per-token discount and the total trajectory saving is your dynamic tail: tool results, retrieved documents, user messages. That portion is paid at full rate no matter what.
A recent arXiv study of caching under agentic workloads makes the same point: although major providers offer prompt caching, its benefits for long-horizon, tool-heavy agents remain underexplored and highly workload-dependent.
Latency and time to first token (TTFT) improvements
Cache hits don't just cut cost; they cut prefill time, which dominates TTFT on long prompts. For an agent making dozens of back-to-back calls, trimming a second or two off each TTFT compounds into noticeably snappier runs.
Prompt caching across OpenAI, Anthropic, and Google
Three providers, three different models for the same underlying mechanism:
- OpenAI automatic caching. On by default for supported models; no API changes required.
- Anthropic explicit cache breakpoints. You mark up to four breakpoints with
cache_control; the SDK handles longest-prefix matching on read. - Google implicit caching + explicit context caching. Implicit is best-effort and free; explicit requires creating a cached-content resource and gives you predictable discounts.
OpenAI: automatic caching and prompt_cache_key
OpenAI's caching is on by default. On GPT-5.5 and earlier, cache hits occur in 128-token increments after the first 1,024 tokens, and a single character difference in that opening 1,024 produces a cache miss with cached_tokens: 0. Supplying a prompt_cache_key lets you steer requests from the same session to the same backend worker and raises your hit rate under load.
Anthropic: explicit cache breakpoints and TTL
Anthropic takes the opposite stance: you mark breakpoints explicitly with cache_control on a content block. Each marked block writes exactly one cache entry hashed to the prefix ending at that block, and automatic prefix checking finds the longest matching prior prefix on reads. Four breakpoints is enough to tier a prompt (tools → system → long context → recent turns) and pick different TTLs where they matter. Anthropic's own guidance recommends caching for conversational agents, coding assistants, and large document workflows, all shapes where a long context gets referenced repeatedly.
Google Gemini: implicit caching vs. explicit context caching
Gemini offers both modes. Implicit caching is automatic and best-effort: no API changes, no guaranteed hit. Explicit context caching requires you to create a cached-content resource up front and reference it; in exchange you get predictable discounts and TTL control. For a long-lived agent with a stable system prompt and tool catalog, explicit context caching is the right default; for ad-hoc traffic, implicit is a free lunch when it works.
Structuring an agent prompt to maximize cache hits
Stable prefix first, dynamic content last
The order inside your prompt is a cost decision. Put everything stable at the top (system prompt, tool definitions, long reference documents, few-shot examples) and push anything that changes per call to the bottom. The ordering that reads naturally in a notebook ("here's today's date, now here's a 2000-token system prompt") is the ordering that will destroy your cache hit rate. The cardinal rule is to design your prompts with stable prefixes and append-only tails.
Keep tool definitions stable and serialization deterministic
Tool schemas are usually the biggest cacheable asset in an agent prompt, and they're the easiest to accidentally invalidate. Non-deterministic JSON serialization (key order changing between runs, trailing whitespace differences, a timestamp or request ID that sneaks into a description field) produces a different hash and a cache miss, even though the schemas are semantically identical. Freeze the serialization. Sort keys. Strip volatile fields.
The same discipline applies to the system prompt. Interpolating datetime.now() into the first paragraph is one of the most common ways teams quietly multiply their input costs, because a refactor that moves a timestamp to the top of the system prompt won't break a single test. It will just show up on next month's invoice.
How dynamic content and history truncation break the cache
Truncating history to fit the context window is necessary, but if you truncate from the front of the message list you invalidate everything downstream. Prefer summarization over head-truncation, and keep the summary in a stable slot. Our agent runtime handles this by proactively compacting old tool results, replacing them with placeholders while preserving the last few user turns, and only falling back to full summarization when the context is still over budget. That keeps the cacheable head intact while bounding the uncacheable tail.
For deeper context on where prompt caching for AI agents fits among the other levers, see our pillar on designing token-efficient AI systems.
Monitoring cache performance in production
Reading the cached_tokens field and provider usage metadata
Every provider reports cache usage on the response:
| Provider | Field | Location |
|---|---|---|
| OpenAI | cached_tokens | usage.prompt_tokens_details |
| Anthropic | cache_creation_input_tokens, cache_read_input_tokens | usage |
| Google Gemini | cached token count | usageMetadata |
Log all of them. The ratio of cache reads to total prompt tokens is your effective cache hit rate, and it's the only honest signal of whether your prefix is actually stable.
Caching fails silently. There is no exception, no warning; a missed cache looks exactly like a successful call, only with a bigger bill. Monitoring the hit rate is the only way to catch a regression before you discover it on an invoice weeks later.
Observability with LangSmith and per-route cost attribution
Aggregate the cached_tokens counts by agent, by tool, and by route so a regression points at a specific change. LangSmith, OpenTelemetry-based tracing, and in-house dashboards all work; what matters is that the hit rate is a tracked metric with an alert, not a number you check manually.
Security and privacy risks of prompt caching
Caching introduces a timing side channel. Because cached prompts are processed faster than uncached ones, an attacker who can measure response latency may be able to infer whether a given prefix is in the cache, and if the cache is shared across users, that can leak information about other users' prompts.
Most providers mitigate this by scoping caches to an organization or an API key, but the exact policy varies and is rarely documented in detail. Don't cache content that contains secrets or user-specific sensitive data in a context where another tenant could probe it, and treat the cache scope as part of your threat model.
Framework and cross-provider considerations
Caching semantics are provider-specific, which makes multi-provider abstractions leaky. A framework that lets you swap OpenAI for Anthropic in one line probably doesn't translate Anthropic's explicit cache breakpoints into OpenAI's prompt_cache_key steering, and the "it still works" result is often "it still works, with the cache off."
If you build on an agent framework, verify what it sends on the wire. If you build on infrastructure, prefer one that treats caching as a first-class concern. We co-locate our agent runtime with retrieval and rerank on one platform, so RAG context stays hot and the system-plus-tools prefix stays byte-stable across the steps of a run, which is what the cache needs to actually fire. Teams that otherwise stitch together a separate vector DB, agent framework, workflow engine, and LLM gateway have more surface area where a stray timestamp or re-serialization can quietly invalidate the prefix; collapsing that stack onto a single platform removes those seams.
The same principle applies in sectors where agent loops run constantly in the background; see our take on the industries most disrupted by agentic AI in 2026 for where this cost math shows up at scale.
Making caching pay off in long-horizon agents
The savings come from one fact: in a long agent loop, the same prefix is sent over and over, and the KV cache lets you pay for it once. Everything else (cache breakpoint placement, serialization determinism, where you truncate, how you summarize) is in service of keeping that prefix exactly identical from step to step.
Three concrete moves to make this week:
- Pin a stable prefix at the top of every agent request and push dynamic content to the bottom.
- Log
cached_tokensandcache_read_input_tokens, and alert when the hit-rate ratio drops. - Audit the system prompt for any per-call value (timestamps, request IDs, non-deterministic JSON) and move it out of the cached region.
Done well, those three changes are the difference between the headline per-token discount and the real double-digit reduction that lands on the invoice.
FAQ
Keep reading
Agents & Workflows Long-Running Agents on Postgres: Durable, Resumable State
How to build long-running agents on Postgres with durable, resumable state, so workflows survive failures and pick up where they left off.
Agents & Workflows Best Automation Tool 2026: viaSocket vs Zapier, Make, n8n & Powabase
Searching for the best automation tool 2026? We compare viaSocket, Zapier, Make, n8n, and Powabase so you can pick the right fit for your workflows.
Agents & Workflows Agent Backend: One Postgres vs. Redis + Kafka + Pinecone
A 200-line Postgres agent orchestrator shows your database can be the framework. What that agent backend gets right, and what production still needs.