CacheCanary

Prompt caching with LangChain on Amazon Bedrock

LangChain talks to Bedrock through the langchain-aws package, and its agents add caching with a middleware. I captured the exact requests langchain-aws 1.8.1 and LangChain 1.4.3 send while create_agent runs its real agent loop, checked them with CacheCanary, and ran every case on live Bedrock (Claude Sonnet 4.6, us-west-2, October 2026).

The short version: caching is off until you add the middleware, and with ChatBedrockConverse it then works well, even with many tool calls at once. Three things quietly throw it away: using ChatBedrock instead, LangChain's own context editing, and adding a cache point of your own.

The setup I'd use

from langchain.agents import create_agent
from langchain_aws import ChatBedrockConverse
from langchain_aws.middleware.prompt_caching import BedrockPromptCachingMiddleware

agent = create_agent(
    model=ChatBedrockConverse(model="us.anthropic.claude-sonnet-4-6"),
    tools=tools,
    system_prompt=SYSTEM_PROMPT,
    middleware=[BedrockPromptCachingMiddleware()],
)

Don't add cache points of your own on top (finding 4), and if you clear old tool results, clear them in batches (finding 3).

What I found

1. Caching is off until you add the middleware

A plain ChatBedrockConverse in an agent sends no cache points. On live Bedrock, the same system prompt in two fresh agents:

SetupSecond call: read from cache
No middleware (the default)0
BedrockPromptCachingMiddleware()3,298

Outside an agent, pass the same setting per call: model.invoke(messages, cache_control={"type": "ephemeral"}). With the middleware, ChatBedrockConverse puts a cache point after the tool list, after the system prompt, at the end of the newest message and at the end of the message before it.

2. Use ChatBedrockConverse, not ChatBedrock

Bedrock only looks back 21 blocks from a cache point to find the previous cache entry, and every tool call and every tool result is a block. With ChatBedrockConverse, the point on the message before the newest one sits right after the model's tool calls, close enough to find the previous entry, so the cache holds for up to 21 tool calls at once. ChatBedrock (the InvokeModel API) gets only one point, at the end of the newest message, and sends the system prompt as plain text with no point of its own. So at 11 tool calls the previous entry is out of reach, and there is nothing earlier to fall back on: the whole prompt, system prompt included, is written to the cache again. On live Bedrock, with a long customer history in the first message:

Next stepRead from cacheWritten again
ChatBedrock, 10 tool calls at once6,9301,783
ChatBedrock, 11 tool calls at once08,890
ChatBedrockConverse, 21 tool calls at once6,9303,708
ChatBedrockConverse, 22 tool calls at once2,222 (system prompt only)8,587

Writing to the cache costs 1.25 times the normal input price, and a cache read a tenth of it, so each of those misses costs 12.5 times more than it should for the part written again.

3. Context editing rewrites the conversation on every call

ContextEditingMiddleware with ClearToolUsesEdit replaces old tool results with "[cleared]" once the conversation passes a token count (100,000 by default), keeping the newest 3. With its default settings, it does this again on every call without saving it, so each new tool result pushes one more old result out of the newest 3. An earlier part of the history changes on every call, and the cache points that come after it can't find the previous entry. Once it starts, every call writes most of the conversation to the cache again. (Setting clear_at_least stops the rewrites, because the same oldest results get cleared every time, but then it never frees more than that many tokens.)

The fix is to clear in batches, so the history only changes once per batch:

from dataclasses import dataclass

from langchain.agents.middleware import ClearToolUsesEdit, ContextEditingMiddleware
from langchain_core.messages import ToolMessage


@dataclass(slots=True)
class ClearToolUsesInChunks(ClearToolUsesEdit):
    """Clear old tool results in whole batches, not one per call."""

    chunk: int = 10

    def apply(self, messages, *, count_tokens) -> None:
        results = sum(isinstance(m, ToolMessage) for m in messages)
        cleared = max(0, results - self.keep) // self.chunk * self.chunk
        keep, self.keep = self.keep, results - cleared  # whole batches only
        try:
            ClearToolUsesEdit.apply(self, messages, count_tokens=count_tokens)
        finally:
            self.keep = keep


editing = ContextEditingMiddleware(edits=[ClearToolUsesInChunks(chunk=10)])
middleware = [editing, BedrockPromptCachingMiddleware()]

On live Bedrock, an agent making one tool call per step with long tool results, with clearing set to start at the seventh step:

ClearingStep 12: read from cacheWritten againCost of step 12
None13,8346272,168
ClearToolUsesEdit (stock)2,2257,2129,238
ClearToolUsesInChunks(chunk=5)11,0446271,889

Cost in full-price input tokens: reads count a tenth, writes 1.25.

So the stock edit, meant to save tokens, cost about 4 times as much per step as not clearing at all, while clearing in batches cost a little less. Clearing in batches rewrote the history on 2 calls out of 31 in my offline run, against every call for the stock edit, and the newest tool results are never cleared. Clearing in batches keeps up to keep + chunk - 1 recent results instead of exactly keep, so the model sees a few more. It also works with exclude_tools and clear_tool_inputs.

4. A cache point of your own turns off LangChain's extra point

Older guides tell you to add ChatBedrockConverse.create_cache_point() to the system message. If you keep that and also use the middleware, LangChain skips the point on the message before the newest one, because it only adds it when the request has no cache points of yours. The limit drops from 21 tool calls at once back to 10. On live Bedrock, 11 tool calls at once:

SetupRead from cacheWritten again
Middleware only6,9341,958
Your own system cache point + middleware2,226 (system prompt only)6,666

The middleware already caches the system prompt, so remove your own cache points when you use it.

Smaller things

What LangChain gets right

The point on the message before the newest one is what lets ChatBedrockConverse handle 21 tool calls at once. The middleware doesn't save cache points into your conversation history or checkpoints, so restored threads stay clean. Structured output with response_format adds its tool to every call, so it doesn't break the cache. It worked with extended thinking turned on, and a cache point in a tool result's content is moved to where Bedrock accepts it.

Check your own agent

Save the exact requests your model sends. This works for ChatBedrockConverse and ChatBedrock, streaming or not; I checked each saved file is identical to what reached Bedrock.

import pathlib, time


def save_requests(model, folder="langchain-requests"):
    """Save every request this LangChain Bedrock model sends, for cachecanary."""
    out = pathlib.Path(folder)
    out.mkdir(exist_ok=True)

    def save(params, **kwargs):
        (out / f"{time.time_ns()}.json").write_bytes(params["body"])

    for op in ("Converse", "ConverseStream",
               "InvokeModel", "InvokeModelWithResponseStream"):
        model.client.meta.events.register(
            f"before-call.bedrock-runtime.{op}", save)


model = ChatBedrockConverse(model="us.anthropic.claude-sonnet-4-6")
save_requests(model)

Run a conversation, then compare two requests in a row with cachecanary diff. These are real outputs on requests LangChain built:

$ cachecanary diff invoke-turn1.json invoke-turn2.json --model us.anthropic.claude-sonnet-4-6
lookback-exceeded: 22 blocks were added between the previous checkpoint and the next one. Bedrock only finds an earlier cache entry up to 21 blocks back (21 added still hits, 22 misses), so this call writes the cache again. Add a checkpoint in between (common after many parallel tool calls). (messages[2].content[10])

$ cachecanary diff edit-call8.json edit-call9.json --model us.anthropic.claude-sonnet-4-6
history-changed: Earlier conversation history was edited (not just appended). (messages[12].content[0])

$ cachecanary diff date-day1.json date-day2.json --model us.anthropic.claude-sonnet-4-6
system-changed: The system prompt changed inside the cached prefix (look for dates, IDs or per-user text). (system[0])

In plain words: the first pair is ChatBedrock with 11 tool calls at once (finding 2). The second is the stock context editing clearing one more tool result (finding 3). The third is a date in a dynamic system prompt.

Keep a few saved requests in your repo and run cachecanary lint on them in CI, and a LangChain upgrade that changes your cache points shows up in a pull request instead of on the bill. Setup is on the home page.

How I tested this

langchain-aws 1.8.1, LangChain 1.4.3 and LangGraph 1.2.14 in a clean Python environment. To capture requests, I pointed LangChain at a small local server standing in for Bedrock that speaks the Converse and InvokeModel formats, streaming or not, so create_agent ran its real agent loop and every request body was saved. The settings cases ran fully live on Bedrock (Claude Sonnet 4.6, us-west-2, October 7 2026). A live model won't call exactly 21 tools on request, so for the agent-loop cases I sent the exact requests LangChain built, in order, to live Bedrock; caching depends only on the request. Every case uses made-up prompts with a fresh random tag, so old cache entries couldn't help. One run read nothing at all from the cache, system prompt included, and four reruns of the same case kept it. That pattern fits cross-Region inference: with a us. model ID, two calls can land in different Regions, and AWS notes this can mean more cache writes. The capture and the live checks are scripts in the repo: capture and live checks. Token counts move by a few tokens from run to run because each run uses a new random tag. LangChain changes quickly, so check your own version with the snippet above.

Same tests for other frameworks: LiteLLM and Strands Agents. More on why caching breaks on Bedrock: Five ways Claude prompt caching quietly breaks on Amazon Bedrock.