CacheCanary

Prompt caching with Strands Agents on Amazon Bedrock

Strands Agents is AWS's open-source agent SDK, and Bedrock is its default model provider. I captured the exact requests Strands 1.58.1 sends to Bedrock while it runs its real agent loop, checked them with CacheCanary, and ran every case on live Bedrock (Claude Sonnet 4.6, us-west-2, October 2026).

The short version: caching is off unless you turn it on. Once it's on, it works for simple calls, but two things that agents do all the time quietly throw the cache away: many tool calls at once, and long conversations.

The setup I'd use

For most agents, turn caching on and swap the conversation manager for one that trims old messages in bulk (finding 3 explains why):

from strands import Agent
from strands.agent.conversation_manager import SlidingWindowConversationManager
from strands.models import BedrockModel
from strands.models.model import CacheConfig


class TrimInChunks(SlidingWindowConversationManager):
    """Trim old messages in bulk, so the cached start rarely changes."""

    def __init__(self, window_size: int = 40, trim_to: int = 20, **kwargs):
        super().__init__(window_size=window_size, **kwargs)
        self.trim_to = trim_to

    def apply_management(self, agent, **kwargs) -> None:
        if len(agent.messages) <= self.window_size:
            return
        full, self.window_size = self.window_size, self.trim_to
        try:
            self.reduce_context(agent)
        finally:
            self.window_size = full

    def restore_from_session(self, state):
        # Let sessions saved with the default manager switch to this one.
        if state.get("__name__") == "SlidingWindowConversationManager":
            state = {**state, "__name__": type(self).__name__}
        return super().restore_from_session(state)


agent = Agent(
    model=BedrockModel(
        model_id="us.anthropic.claude-sonnet-4-6",
        cache_config=CacheConfig(strategy="auto"),
    ),
    system_prompt=SYSTEM_PROMPT,
    tools=tools,
    conversation_manager=TrimInChunks(),
)

If your agent sometimes calls more than 10 tools at once, add the hook from finding 2 as well.

What I found

1. Caching is off by default

A plain BedrockModel(model_id=...) sends no cache points, so every call pays full price for the whole prompt. A change to turn caching on by default was proposed and closed without being merged. On live Bedrock, the same system prompt sent twice:

SetupSecond call: read from cachePaid at full price
No cache_config (the default)01,996
CacheConfig(strategy="auto")1,9933

With strategy="auto", Strands puts one cache point after the system prompt and one at the end of the newest user message. After a round of tool calls, the newest user message is the one holding the tool results.

2. More than 10 tool calls at once lose the conversation

Bedrock only looks back 21 blocks from a cache point to find the previous cache entry, and every tool call and every tool result is a block. Auto mode keeps a single cache point in the conversation and moves it to the newest message, so a step that runs 11 tools at once adds 22 blocks and the previous entry is out of reach. The whole conversation is written to the cache again, at 1.25 times the normal price instead of a tenth of it. On live Bedrock, with a long customer history in the first message:

Next stepRead from cacheWritten again
10 tool calls at once6,9331,783
11 tool calls at once2,221 (system prompt only)6,662

This is issue #3348, still open. Auto mode removes any cache point you add to earlier messages, so the fix is to leave cache_config unset and place the points yourself with a hook. This one uses the 4 points Bedrock allows on the system prompt, your newest message, the first result of the newest round of tool calls, and the end of the conversation:

from strands.hooks import BeforeModelCallEvent, HookProvider, HookRegistry


class BedrockCachePoints(HookProvider):
    """Place cache points in the messages before every model call."""

    def register_hooks(self, registry: HookRegistry, **kwargs) -> None:
        registry.add_callback(BeforeModelCallEvent, self.place)

    def place(self, event: BeforeModelCallEvent) -> None:
        messages = event.agent.messages
        for message in messages:  # drop the points from the previous call
            message["content"] = [b for b in message["content"]
                                  if "cachePoint" not in b]
        typed = [i for i, m in enumerate(messages) if m["role"] == "user"
                 and not any("toolResult" in b for b in m["content"])]
        if typed:  # the newest message a person typed
            newest = messages[typed[-1]]["content"]
            self._add(newest, len(newest))
        last = messages[-1]["content"] if messages else []
        if messages and (not typed or typed[-1] != len(messages) - 1):
            if any("toolResult" in b for b in last) and len(last) > 1:
                self._add(last, 1)  # after the first result of the round
            self._add(last, len(last))  # the end of the conversation

    @staticmethod
    def _add(content, index):
        # Bedrock rejects a cache point right after a non-PDF document,
        # so step back past those.
        def other_document(block):
            return block.get("document", {}).get("format", "pdf") != "pdf"

        while index > 0 and other_document(content[index - 1]):
            index -= 1
        if index > 0 and "cachePoint" not in content[index - 1]:
            content.insert(index, {"cachePoint": {"type": "default"}})


agent = Agent(
    # no cache_config: the hook places the message cache points
    model=BedrockModel(model_id="us.anthropic.claude-sonnet-4-6"),
    system_prompt=[{"text": SYSTEM_PROMPT}, {"cachePoint": {"type": "default"}}],
    tools=tools,
    hooks=[BedrockCachePoints()],
    conversation_manager=TrimInChunks(),
)

On live Bedrock, through Strands' own agent loop, each step read almost the whole previous request from the cache:

StepRead from cacheSize of the previous request
12 tool calls at once6,9336,936
1 tool call9,0669,067
20 tool calls at once9,2459,246
A follow-up question from the user12,77512,776
3 tool calls at once12,79112,794

Tokens. The last few tokens of each request are always new, so the cache can't hold them.

It holds for up to 20 tool calls at once. At 21, the round before it is written again, but the point on your newest message still holds everything up to there (12,791 of 13,350 tokens read). I also ran it with a live model choosing its own tools: it called 6 at once, then 2, and the second question read 7,108 tokens from the cache and wrote 485. It works with Strands' session managers (they don't save the cache points, so a restored session starts clean) and with extended thinking turned on.

3. Long conversations lose the cache on every turn

By default Strands keeps the last 40 messages. Once a conversation is longer, it drops the oldest messages at the end of every call to the agent, so the start of the conversation changes every turn and nothing after the system prompt can be read from the cache. On live Bedrock, a chat with full-length answers, from the 22nd question on:

Conversation managerRead from cacheWritten again
Default (40 messages)2,225 on every turn15,871 on every turn
TrimInChunks()10,323, then 11,100, 11,877, 12,654777 on each turn

TrimInChunks (in the setup above) lets the history grow to 40 messages and then trims it back to 20 in one go, so the start of the conversation changes about once every 10 turns instead of every turn. Over 45 questions, the default manager changed it on 24 calls and TrimInChunks on 3. It keeps 20 to 40 messages instead of always 40, so the model sees a little less history; raise both numbers if you need more. Without tools or with them, it works the same, and it works with the hook from finding 2 and with per_turn and pin_first. The last method matters if you save sessions: Strands checks the manager's class name when it restores one, and without it, a session saved with the default manager fails to load ("Invalid conversation manager state").

4. Application inference profiles need strategy="anthropic"

Auto mode only caches when the model ID contains "claude" or "anthropic". An application inference profile ARN contains neither, so Strands logs a warning ("cache_config is enabled but this model does not support automatic caching") and sends no cache points. The maintainers chose this on purpose. If the profile points to a Claude model, say so:

BedrockModel(
    model_id="arn:aws:bedrock:REGION:ACCOUNT:application-inference-profile/ID",
    cache_config=CacheConfig(strategy="anthropic"),
)

On live Bedrock with a profile for Sonnet 4.6, the second call read 0 tokens from the cache with "auto" and 1,991 with "anthropic".

5. Two cache lifetime settings that Bedrock rejects

Bedrock requires a 1-hour cache point to come before any 5-minute one, in the order tools, system prompt, messages. CacheConfig(ttl="1h", system_prompt_ttl="5m") and CacheConfig(ttl="1h", tools_ttl="5m") both put a 5-minute point first, so Bedrock rejects every request with "a ttl='1h' cache_control block must not come after a ttl='5m' cache_control block". This one is loud, but it takes your agent down until you change it (issue #3758). Use one lifetime everywhere, or give the earlier sections the longer one.

Smaller things

What Strands gets right

In auto mode Strands caches the system prompt as well as the conversation, keeps a cache point you place in the newest user message, and steps a cache point back when it would land right after a non-PDF document, which Bedrock rejects. In my tests, every request it sent was valid for Bedrock, apart from the two lifetime settings above.

Check your own agent

Save the exact requests your BedrockModel sends. This works for streaming and non-streaming calls; I checked each saved file is identical to what reached Bedrock.

import pathlib, time


def save_requests(model, folder="strands-requests"):
    """Save every request this BedrockModel sends, for cachecanary."""
    out = pathlib.Path(folder)
    out.mkdir(exist_ok=True)

    def save(params, **kwargs):
        (out / f"{time.time_ns()}.json").write_bytes(params["body"])

    for op in ("Converse", "ConverseStream"):
        model.client.meta.events.register(
            f"before-call.bedrock-runtime.{op}", save)


model = BedrockModel(
    model_id="us.anthropic.claude-sonnet-4-6",
    cache_config=CacheConfig(strategy="auto"),
)
save_requests(model)

Run a conversation, then compare two requests in a row with cachecanary diff, or check one with cachecanary lint. These are real outputs on requests Strands built:

$ cachecanary diff agent-turn1.json agent-turn2.json --model us.anthropic.claude-sonnet-4-6
lookback-exceeded: 22 blocks were added between the previous checkpoint and the next one. Bedrock only finds an earlier cache entry up to 21 blocks back (21 added still hits, 22 misses), so this call writes the cache again. Add a checkpoint in between (common after many parallel tool calls). (messages[2].content[10])

$ cachecanary diff chat-turn20.json chat-turn21.json --model us.anthropic.claude-sonnet-4-6
history-changed: Earlier conversation history was edited (not just appended). (messages[0].content[0])

$ cachecanary lint ttl-order.json --model us.anthropic.claude-sonnet-4-6
[error] ttl-order: A 1h checkpoint appears after a 5m checkpoint. Longer TTLs must come first.

In plain words: the first pair is the 11-tool-call step from finding 2. The second is the default conversation manager trimming the first messages (finding 3). The third is the lifetime setting Bedrock rejects (finding 5).

Keep a few saved requests in your repo and run cachecanary lint on them in CI, and a Strands upgrade that changes your cache points shows up in a pull request instead of on the bill. Setup is on the home page.

How I tested this

Strands Agents 1.58.1 in a clean Python environment. To capture requests, I pointed Strands at a small local server standing in for Bedrock that speaks its streaming format, so Strands ran its real agent loop (tool calls, streaming, conversation trimming) and every request body was saved. The settings cases ran fully live on Bedrock (Claude Sonnet 4.6, us-west-2, October 7 2026). A live model won't call exactly 11 tools on request, so for the agent-loop cases I sent the exact requests Strands built, in order, to live Bedrock; caching depends only on the request. Every case uses made-up prompts with a fresh random tag, so old cache entries couldn't help. The capture and the live checks are scripts in the repo, so you can run them yourself: capture and live checks. Token counts move by a few tokens from run to run because each run uses a new random tag. Strands changes quickly, so check your own version with the snippet above.

Using LiteLLM instead? Same tests: Prompt caching with LiteLLM on Amazon Bedrock. More on why caching breaks on Bedrock: Five ways Claude prompt caching quietly breaks on Amazon Bedrock.