Prompt caching with LangChain on Amazon Bedrock
LangChain talks to Bedrock through the langchain-aws package, and its agents add caching with a middleware. I captured the exact requests langchain-aws 1.8.1 and LangChain 1.4.3 send while create_agent runs its real agent loop, checked them with CacheCanary, and ran every case on live Bedrock (Claude Sonnet 4.6, us-west-2, October 2026).
The short version: caching is off until you add the middleware, and with ChatBedrockConverse it then works well, even with many tool calls at once. Three things quietly throw it away: using ChatBedrock instead, LangChain's own context editing, and adding a cache point of your own.
The setup I'd use
from langchain.agents import create_agent
from langchain_aws import ChatBedrockConverse
from langchain_aws.middleware.prompt_caching import BedrockPromptCachingMiddleware
agent = create_agent(
model=ChatBedrockConverse(model="us.anthropic.claude-sonnet-4-6"),
tools=tools,
system_prompt=SYSTEM_PROMPT,
middleware=[BedrockPromptCachingMiddleware()],
)
Don't add cache points of your own on top (finding 4), and if you clear old tool results, clear them in batches (finding 3).
What I found
1. Caching is off until you add the middleware
A plain ChatBedrockConverse in an agent sends no cache points. On live Bedrock, the same system prompt in two fresh agents:
| Setup | Second call: read from cache |
|---|---|
| No middleware (the default) | 0 |
BedrockPromptCachingMiddleware() | 3,298 |
Outside an agent, pass the same setting per call: model.invoke(messages, cache_control={"type": "ephemeral"}). With the middleware, ChatBedrockConverse puts a cache point after the tool list, after the system prompt, at the end of the newest message and at the end of the message before it.
2. Use ChatBedrockConverse, not ChatBedrock
Bedrock only looks back 21 blocks from a cache point to find the previous cache entry, and every tool call and every tool result is a block. With ChatBedrockConverse, the point on the message before the newest one sits right after the model's tool calls, close enough to find the previous entry, so the cache holds for up to 21 tool calls at once. ChatBedrock (the InvokeModel API) gets only one point, at the end of the newest message, and sends the system prompt as plain text with no point of its own. So at 11 tool calls the previous entry is out of reach, and there is nothing earlier to fall back on: the whole prompt, system prompt included, is written to the cache again. On live Bedrock, with a long customer history in the first message:
| Next step | Read from cache | Written again |
|---|---|---|
ChatBedrock, 10 tool calls at once | 6,930 | 1,783 |
ChatBedrock, 11 tool calls at once | 0 | 8,890 |
ChatBedrockConverse, 21 tool calls at once | 6,930 | 3,708 |
ChatBedrockConverse, 22 tool calls at once | 2,222 (system prompt only) | 8,587 |
Writing to the cache costs 1.25 times the normal input price, and a cache read a tenth of it, so each of those misses costs 12.5 times more than it should for the part written again.
3. Context editing rewrites the conversation on every call
ContextEditingMiddleware with ClearToolUsesEdit replaces old tool results with "[cleared]" once the conversation passes a token count (100,000 by default), keeping the newest 3. With its default settings, it does this again on every call without saving it, so each new tool result pushes one more old result out of the newest 3. An earlier part of the history changes on every call, and the cache points that come after it can't find the previous entry. Once it starts, every call writes most of the conversation to the cache again. (Setting clear_at_least stops the rewrites, because the same oldest results get cleared every time, but then it never frees more than that many tokens.)
The fix is to clear in batches, so the history only changes once per batch:
from dataclasses import dataclass
from langchain.agents.middleware import ClearToolUsesEdit, ContextEditingMiddleware
from langchain_core.messages import ToolMessage
@dataclass(slots=True)
class ClearToolUsesInChunks(ClearToolUsesEdit):
"""Clear old tool results in whole batches, not one per call."""
chunk: int = 10
def apply(self, messages, *, count_tokens) -> None:
results = sum(isinstance(m, ToolMessage) for m in messages)
cleared = max(0, results - self.keep) // self.chunk * self.chunk
keep, self.keep = self.keep, results - cleared # whole batches only
try:
ClearToolUsesEdit.apply(self, messages, count_tokens=count_tokens)
finally:
self.keep = keep
editing = ContextEditingMiddleware(edits=[ClearToolUsesInChunks(chunk=10)])
middleware = [editing, BedrockPromptCachingMiddleware()]
On live Bedrock, an agent making one tool call per step with long tool results, with clearing set to start at the seventh step:
| Clearing | Step 12: read from cache | Written again | Cost of step 12 |
|---|---|---|---|
| None | 13,834 | 627 | 2,168 |
ClearToolUsesEdit (stock) | 2,225 | 7,212 | 9,238 |
ClearToolUsesInChunks(chunk=5) | 11,044 | 627 | 1,889 |
Cost in full-price input tokens: reads count a tenth, writes 1.25.
So the stock edit, meant to save tokens, cost about 4 times as much per step as not clearing at all, while clearing in batches cost a little less. Clearing in batches rewrote the history on 2 calls out of 31 in my offline run, against every call for the stock edit, and the newest tool results are never cleared. Clearing in batches keeps up to keep + chunk - 1 recent results instead of exactly keep, so the model sees a few more. It also works with exclude_tools and clear_tool_inputs.
4. A cache point of your own turns off LangChain's extra point
Older guides tell you to add ChatBedrockConverse.create_cache_point() to the system message. If you keep that and also use the middleware, LangChain skips the point on the message before the newest one, because it only adds it when the request has no cache points of yours. The limit drops from 21 tool calls at once back to 10. On live Bedrock, 11 tool calls at once:
| Setup | Read from cache | Written again |
|---|---|---|
| Middleware only | 6,934 | 1,958 |
| Your own system cache point + middleware | 2,226 (system prompt only) | 6,666 |
The middleware already caches the system prompt, so remove your own cache points when you use it.
Smaller things
SummarizationMiddlewarereplaces old messages with a summary whenever it runs, which rewrites the conversation once each time. On live Bedrock, each call right after a summary read only the system prompt from the cache, and the calls after it read normally again. The gap between itstriggerandkeepsettings decides how often this happens; keep it wide.- The older
create_react_agentfrom LangGraph (deprecated in favour ofcreate_agent) doesn't run middleware, andmodel.bind(cache_control=...)doesn't reach the request through it either: I got no cache points at all. A cache point in the system message caches only the system prompt. Move tocreate_agentwith the middleware. - With an application inference profile ARN, LangChain needs
provider="anthropic"(it raises an error without it). It then looks up the model behind the profile with aGetInferenceProfilecall when the model is created, which needs that IAM permission, or you can passbase_model_id="anthropic.claude-sonnet-4-6". Both cached on live Bedrock. - If you add up cache writes from
usage_metadata, readephemeral_5m_input_tokensandephemeral_1h_input_tokens. When Bedrock reports the cache lifetime, LangChain puts the writes there and setscache_creationto 0, on purpose. - Claude Haiku 4.5 needs at least 4,096 tokens before anything is cached. A 3,850-token prompt cached nothing on live Bedrock, with no error.
- Anything that changes in a
@dynamic_promptsystem prompt starts a new cache: a date once a day, a time or a request ID on every call.cachecanary lintwarns about dates, times, IDs and timestamps.
What LangChain gets right
The point on the message before the newest one is what lets ChatBedrockConverse handle 21 tool calls at once. The middleware doesn't save cache points into your conversation history or checkpoints, so restored threads stay clean. Structured output with response_format adds its tool to every call, so it doesn't break the cache. It worked with extended thinking turned on, and a cache point in a tool result's content is moved to where Bedrock accepts it.
Check your own agent
Save the exact requests your model sends. This works for ChatBedrockConverse and ChatBedrock, streaming or not; I checked each saved file is identical to what reached Bedrock.
import pathlib, time
def save_requests(model, folder="langchain-requests"):
"""Save every request this LangChain Bedrock model sends, for cachecanary."""
out = pathlib.Path(folder)
out.mkdir(exist_ok=True)
def save(params, **kwargs):
(out / f"{time.time_ns()}.json").write_bytes(params["body"])
for op in ("Converse", "ConverseStream",
"InvokeModel", "InvokeModelWithResponseStream"):
model.client.meta.events.register(
f"before-call.bedrock-runtime.{op}", save)
model = ChatBedrockConverse(model="us.anthropic.claude-sonnet-4-6")
save_requests(model)
Run a conversation, then compare two requests in a row with cachecanary diff. These are real outputs on requests LangChain built:
$ cachecanary diff invoke-turn1.json invoke-turn2.json --model us.anthropic.claude-sonnet-4-6 lookback-exceeded: 22 blocks were added between the previous checkpoint and the next one. Bedrock only finds an earlier cache entry up to 21 blocks back (21 added still hits, 22 misses), so this call writes the cache again. Add a checkpoint in between (common after many parallel tool calls). (messages[2].content[10]) $ cachecanary diff edit-call8.json edit-call9.json --model us.anthropic.claude-sonnet-4-6 history-changed: Earlier conversation history was edited (not just appended). (messages[12].content[0]) $ cachecanary diff date-day1.json date-day2.json --model us.anthropic.claude-sonnet-4-6 system-changed: The system prompt changed inside the cached prefix (look for dates, IDs or per-user text). (system[0])
In plain words: the first pair is ChatBedrock with 11 tool calls at once (finding 2). The second is the stock context editing clearing one more tool result (finding 3). The third is a date in a dynamic system prompt.
Keep a few saved requests in your repo and run cachecanary lint on them in CI, and a LangChain upgrade that changes your cache points shows up in a pull request instead of on the bill. Setup is on the home page.
How I tested this
langchain-aws 1.8.1, LangChain 1.4.3 and LangGraph 1.2.14 in a clean Python environment. To capture requests, I pointed LangChain at a small local server standing in for Bedrock that speaks the Converse and InvokeModel formats, streaming or not, so create_agent ran its real agent loop and every request body was saved. The settings cases ran fully live on Bedrock (Claude Sonnet 4.6, us-west-2, October 7 2026). A live model won't call exactly 21 tools on request, so for the agent-loop cases I sent the exact requests LangChain built, in order, to live Bedrock; caching depends only on the request. Every case uses made-up prompts with a fresh random tag, so old cache entries couldn't help. One run read nothing at all from the cache, system prompt included, and four reruns of the same case kept it. That pattern fits cross-Region inference: with a us. model ID, two calls can land in different Regions, and AWS notes this can mean more cache writes. The capture and the live checks are scripts in the repo: capture and live checks. Token counts move by a few tokens from run to run because each run uses a new random tag. LangChain changes quickly, so check your own version with the snippet above.
Same tests for other frameworks: LiteLLM and Strands Agents. More on why caching breaks on Bedrock: Five ways Claude prompt caching quietly breaks on Amazon Bedrock.