CacheCanary

Prompt caching with LiteLLM on Amazon Bedrock

LiteLLM can add Bedrock cache points for you. I wanted to know what it really sends, so I captured the exact requests LiteLLM 1.104.0 builds for Bedrock, checked them with CacheCanary, and ran the important cases on live Bedrock (Claude Sonnet 4.6, us-west-2, October 2026).

The short version: LiteLLM puts cache points in valid places. The money is lost in where you ask it to put them, and in a few settings it drops without telling you.

The setup I'd use

Bedrock lets you use 4 cache points per request. This helper spends them on the four places that matter: the system prompt, your newest user message, the start of the newest round of tool results, and the end of the conversation.

def cache_points(messages):
    """Bedrock cache points for this call (Bedrock allows 4)."""
    # The system prompt and the end of the conversation.
    points = [{"location": "message", "role": "system"},
              {"location": "message", "index": -1}]
    users = [i for i, m in enumerate(messages) if m["role"] == "user"]
    rounds = [i for i, m in enumerate(messages)
              if m["role"] == "assistant" and m.get("tool_calls")]
    # The newest user message.
    if users:
        points.append({"location": "message", "index": users[-1]})
    # The first tool result of the newest round of tool calls.
    if rounds and (not users or rounds[-1] > users[-1]) \
            and rounds[-1] + 1 < len(messages):
        points.append({"location": "message", "index": rounds[-1] + 1})
    return points


response = litellm.completion(
    model="bedrock/converse/us.anthropic.claude-sonnet-4-6",
    messages=messages,
    tools=tools,
    cache_control_injection_points=cache_points(messages),
)

Work out the points again on every call, because the indexes move as the conversation grows. I ran a 7-step agent session with it on live Bedrock: each step read almost the whole previous prompt from the cache.

StepRead from cacheSize of the previous request
12 parallel tool calls6,9266,929
1 tool call8,8678,868
20 parallel tool calls9,0309,031
A follow-up question from the user12,24012,241
3 parallel tool calls12,25612,259

Tokens. The last few tokens of each request are always new, so the cache can't hold them.

If you use the LiteLLM Proxy, the same points go in the request body. With the OpenAI client:

response = client.chat.completions.create(
    model="claude-sonnet",
    messages=messages,
    tools=tools,
    extra_body={"cache_control_injection_points": cache_points(messages)},
)

What I found

1. role: "user" caches your oldest messages, not your newest

A common setup is one point on the system prompt and one on role: "user". That marks every user message. LiteLLM stops at 4 points, in the order it finds them, so once a chat has more than three user messages, the points sit on the first three and stay there. Everything after them is paid at full price on every call, and that part grows each turn. LiteLLM logs a warning ("Reached the provider limit of 4 cache breakpoints") and sends the request anyway.

On live Bedrock, turn 10 of a chat:

PointsRead from cachePaid at full price
system + role: "user"4,3137,749
system + index: -111,2473

Use index: -1 (the last message) instead of role: "user".

2. Many parallel tool calls lose the conversation

With points on the system prompt and the last message, the conversation point moves to the newest tool result on each step. That works until one step adds a lot of blocks. Bedrock only looks back 21 blocks from a cache point for the previous cache entry, and every tool call and every tool result is a block. 12 parallel calls add 24 blocks, so the previous entry is out of reach and the conversation is written to the cache again. On live Bedrock, with a long customer history in the first user message:

Next stepRead from cacheWritten again
4 parallel tool calls6,920669
12 parallel tool calls2,214 (system prompt only)6,647

The part that was written again costs 1.25 times the normal input price instead of a tenth of it, so 12.5 times more than a cache read. In a long agent session that happens on every wide step. The helper above fixes it with the point on the first result of the newest round: that point is close enough to find the previous entry. It holds for up to 20 parallel calls in one round. At 21, the round before it is written again, but the point on the newest user message still holds everything up to there (in my run: 12,256 of the previous 12,767 tokens read from the cache).

3. A cache point on an assistant tool-call message is dropped

The obvious place for an extra point is the assistant message that holds the tool calls. When that message has only tool calls and no text, LiteLLM has nothing to attach the point to and sends the request without it. No warning. That's why the helper uses the first tool result instead, which LiteLLM turns into a valid Bedrock cache point.

4. The 1-hour cache lifetime can be dropped without a warning

You can ask for a 1-hour cache with "ttl": "1h" in cache_control, or in an injection point's control. LiteLLM only sends it for models whose entry in its price list says they support it. Otherwise it sends a normal 5-minute cache point. I found two cases where that happens to a model that does support 1 hour:

I checked both on live Bedrock. Bedrock's response says which lifetime it used, and in both cases it was 5 minutes. With the fix above (or the downloaded price list), it was 1 hour. The request still works either way, so the only sign is more cache misses after 5 quiet minutes. To check, look for "ttl": "1h" in the request LiteLLM sends (see below).

Smaller things

What LiteLLM gets right

Every cache point LiteLLM sent was in a place Bedrock accepts. A point on a tool result goes after the tool result, not inside it (inside, Bedrock ignores it). The InvokeModel route (bedrock/invoke/) keeps cache_control as it is. The LiteLLM Proxy behaves the same: I ran it with points in config.yaml, with cache_control sent by the client, and with points sent per request, and all three reached Bedrock.

Check your own app

Save the exact requests LiteLLM sends to Bedrock. This callback writes each one to a litellm-requests folder. It works for normal, async and streaming calls; I checked each saved file is identical to what reached Bedrock.

import json, pathlib, time
import litellm
from litellm.integrations.custom_logger import CustomLogger

class SaveBedrockRequests(CustomLogger):
    def log_pre_api_call(self, model, messages, kwargs):
        body = (kwargs.get("additional_args") or {}).get("complete_input_dict")
        if not body:
            return
        if isinstance(body, bytes):
            body = body.decode()
        if not isinstance(body, str):
            body = json.dumps(body, default=str)
        folder = pathlib.Path("litellm-requests")
        folder.mkdir(exist_ok=True)
        (folder / f"{time.time_ns()}.json").write_text(body)

litellm.callbacks = [SaveBedrockRequests()]

Run one conversation, then compare two requests in a row with cachecanary diff. These are real outputs on requests LiteLLM built:

$ cachecanary diff agent-turn1.json agent-turn2.json --model us.anthropic.claude-sonnet-4-6
lookback-exceeded: 24 blocks were added between the previous checkpoint and the next one. Bedrock only finds an earlier cache entry up to 21 blocks back (21 added still hits, 22 misses), so this call writes the cache again. Add a checkpoint in between (common after many parallel tool calls). (messages[2].content[11])

$ cachecanary diff chat-turn6.json chat-turn7.json --model us.anthropic.claude-sonnet-4-6
prefix-identical: The cached prefix is identical. If B still missed, the entry likely expired (TTL elapsed between calls) or cross-region inference routed to a Region without the entry.
checkpoint-not-moved: The cached part was read, but the last cache point didn't move while 8 new blocks were added after it. Those are paid at full price on every call. Put the cache point at the end of the newest message. (messages[4].content[0])

$ cachecanary diff with-download.json without-download.json --model us.anthropic.claude-sonnet-5-5
ttl-changed: Checkpoint TTL changed (1h -> 5m); entries are keyed by TTL.

In plain words: the first pair is the 12-tool-call step from finding 2. The second is the role: "user" chat from finding 1: the cache works, but the newest messages are never added to it. The third is the same Sonnet 5.5 request with and without LiteLLM's price list download (finding 4).

Keep a few saved requests in your repo and run cachecanary lint on them in CI, and a LiteLLM upgrade that changes your cache points shows up in a pull request instead of on the bill. Setup is on the home page.

How I tested this

LiteLLM 1.104.0 in a clean Python environment. To capture requests, I pointed LiteLLM at a small local server standing in for Bedrock that saved every request body. For the live numbers I sent the same requests to Bedrock (Claude Sonnet 4.6, us-west-2, October 7 2026) with made-up prompts and a fresh random tag in each run, so old cache entries couldn't help. The capture and the live checks are scripts in the repo, so you can run them yourself: capture and live checks. I also ran LiteLLM's own proxy server against live Bedrock, called with the OpenAI client as shown above, and got the same numbers as the library. Token counts move by a few tokens from run to run because each run uses a new random tag. LiteLLM changes quickly, so check your own version with the callback above.

More on why caching breaks on Bedrock, with the bug reports behind each case: Five ways Claude prompt caching quietly breaks on Amazon Bedrock.