Prompt caching with LiteLLM on Amazon Bedrock
LiteLLM can add Bedrock cache points for you. I wanted to know what it really sends, so I captured the exact requests LiteLLM 1.104.0 builds for Bedrock, checked them with CacheCanary, and ran the important cases on live Bedrock (Claude Sonnet 4.6, us-west-2, October 2026).
The short version: LiteLLM puts cache points in valid places. The money is lost in where you ask it to put them, and in a few settings it drops without telling you.
The setup I'd use
Bedrock lets you use 4 cache points per request. This helper spends them on the four places that matter: the system prompt, your newest user message, the start of the newest round of tool results, and the end of the conversation.
def cache_points(messages):
"""Bedrock cache points for this call (Bedrock allows 4)."""
# The system prompt and the end of the conversation.
points = [{"location": "message", "role": "system"},
{"location": "message", "index": -1}]
users = [i for i, m in enumerate(messages) if m["role"] == "user"]
rounds = [i for i, m in enumerate(messages)
if m["role"] == "assistant" and m.get("tool_calls")]
# The newest user message.
if users:
points.append({"location": "message", "index": users[-1]})
# The first tool result of the newest round of tool calls.
if rounds and (not users or rounds[-1] > users[-1]) \
and rounds[-1] + 1 < len(messages):
points.append({"location": "message", "index": rounds[-1] + 1})
return points
response = litellm.completion(
model="bedrock/converse/us.anthropic.claude-sonnet-4-6",
messages=messages,
tools=tools,
cache_control_injection_points=cache_points(messages),
)
Work out the points again on every call, because the indexes move as the conversation grows. I ran a 7-step agent session with it on live Bedrock: each step read almost the whole previous prompt from the cache.
| Step | Read from cache | Size of the previous request |
|---|---|---|
| 12 parallel tool calls | 6,926 | 6,929 |
| 1 tool call | 8,867 | 8,868 |
| 20 parallel tool calls | 9,030 | 9,031 |
| A follow-up question from the user | 12,240 | 12,241 |
| 3 parallel tool calls | 12,256 | 12,259 |
Tokens. The last few tokens of each request are always new, so the cache can't hold them.
If you use the LiteLLM Proxy, the same points go in the request body. With the OpenAI client:
response = client.chat.completions.create(
model="claude-sonnet",
messages=messages,
tools=tools,
extra_body={"cache_control_injection_points": cache_points(messages)},
)
What I found
1. role: "user" caches your oldest messages, not your newest
A common setup is one point on the system prompt and one on role: "user". That marks every user message. LiteLLM stops at 4 points, in the order it finds them, so once a chat has more than three user messages, the points sit on the first three and stay there. Everything after them is paid at full price on every call, and that part grows each turn. LiteLLM logs a warning ("Reached the provider limit of 4 cache breakpoints") and sends the request anyway.
On live Bedrock, turn 10 of a chat:
| Points | Read from cache | Paid at full price |
|---|---|---|
system + role: "user" | 4,313 | 7,749 |
system + index: -1 | 11,247 | 3 |
Use index: -1 (the last message) instead of role: "user".
2. Many parallel tool calls lose the conversation
With points on the system prompt and the last message, the conversation point moves to the newest tool result on each step. That works until one step adds a lot of blocks. Bedrock only looks back 21 blocks from a cache point for the previous cache entry, and every tool call and every tool result is a block. 12 parallel calls add 24 blocks, so the previous entry is out of reach and the conversation is written to the cache again. On live Bedrock, with a long customer history in the first user message:
| Next step | Read from cache | Written again |
|---|---|---|
| 4 parallel tool calls | 6,920 | 669 |
| 12 parallel tool calls | 2,214 (system prompt only) | 6,647 |
The part that was written again costs 1.25 times the normal input price instead of a tenth of it, so 12.5 times more than a cache read. In a long agent session that happens on every wide step. The helper above fixes it with the point on the first result of the newest round: that point is close enough to find the previous entry. It holds for up to 20 parallel calls in one round. At 21, the round before it is written again, but the point on the newest user message still holds everything up to there (in my run: 12,256 of the previous 12,767 tokens read from the cache).
3. A cache point on an assistant tool-call message is dropped
The obvious place for an extra point is the assistant message that holds the tool calls. When that message has only tool calls and no text, LiteLLM has nothing to attach the point to and sends the request without it. No warning. That's why the helper uses the first tool result instead, which LiteLLM turns into a valid Bedrock cache point.
4. The 1-hour cache lifetime can be dropped without a warning
You can ask for a 1-hour cache with "ttl": "1h" in cache_control, or in an injection point's control. LiteLLM only sends it for models whose entry in its price list says they support it. Otherwise it sends a normal 5-minute cache point. I found two cases where that happens to a model that does support 1 hour:
- Application inference profiles. When the model is a profile ARN (
bedrock/converse/arn:aws:bedrock:...), LiteLLM can't tell which model it is, so it drops the 1-hour setting. Name the model and pass the ARN separately, and it's kept:# PROFILE_ARN = "arn:aws:bedrock:REGION:ACCOUNT:application-inference-profile/ID" litellm.completion( model="bedrock/converse/us.anthropic.claude-sonnet-4-6", model_id=PROFILE_ARN, ... )On the InvokeModel route (bedrock/invoke/), a profile ARN as the model fails with an error, somodel_idis the way to use one there too. - Claude Sonnet 5.5 with LiteLLM's built-in price list. LiteLLM downloads its price list when it starts. If you set
LITELLM_LOCAL_MODEL_COST_MAP=True, or the download fails (it falls back with a "Falling back to local backup" warning), it uses the copy inside the package, and 1.104.0's copy has no Sonnet 5.5. Sonnet 4.6 and Opus 5.5 are in it and keep 1 hour.
I checked both on live Bedrock. Bedrock's response says which lifetime it used, and in both cases it was 5 minutes. With the fix above (or the downloaded price list), it was 1 hour. The request still works either way, so the only sign is more cache misses after 5 quiet minutes. To check, look for "ttl": "1h" in the request LiteLLM sends (see below).
Smaller things
location: "tool_config"puts a point after your tool list. Bedrock caches tools first, so if the tool list alone is shorter than the model's minimum (1,024 tokens on Sonnet 4.6), that point caches nothing and uses up one of your 4.- A 1-hour point after a 5-minute point is rejected by Bedrock with an error. At least that one is loud.
cachecanary lintcatches it before it ships.
What LiteLLM gets right
Every cache point LiteLLM sent was in a place Bedrock accepts. A point on a tool result goes after the tool result, not inside it (inside, Bedrock ignores it). The InvokeModel route (bedrock/invoke/) keeps cache_control as it is. The LiteLLM Proxy behaves the same: I ran it with points in config.yaml, with cache_control sent by the client, and with points sent per request, and all three reached Bedrock.
Check your own app
Save the exact requests LiteLLM sends to Bedrock. This callback writes each one to a litellm-requests folder. It works for normal, async and streaming calls; I checked each saved file is identical to what reached Bedrock.
import json, pathlib, time
import litellm
from litellm.integrations.custom_logger import CustomLogger
class SaveBedrockRequests(CustomLogger):
def log_pre_api_call(self, model, messages, kwargs):
body = (kwargs.get("additional_args") or {}).get("complete_input_dict")
if not body:
return
if isinstance(body, bytes):
body = body.decode()
if not isinstance(body, str):
body = json.dumps(body, default=str)
folder = pathlib.Path("litellm-requests")
folder.mkdir(exist_ok=True)
(folder / f"{time.time_ns()}.json").write_text(body)
litellm.callbacks = [SaveBedrockRequests()]
Run one conversation, then compare two requests in a row with cachecanary diff. These are real outputs on requests LiteLLM built:
$ cachecanary diff agent-turn1.json agent-turn2.json --model us.anthropic.claude-sonnet-4-6 lookback-exceeded: 24 blocks were added between the previous checkpoint and the next one. Bedrock only finds an earlier cache entry up to 21 blocks back (21 added still hits, 22 misses), so this call writes the cache again. Add a checkpoint in between (common after many parallel tool calls). (messages[2].content[11]) $ cachecanary diff chat-turn6.json chat-turn7.json --model us.anthropic.claude-sonnet-4-6 prefix-identical: The cached prefix is identical. If B still missed, the entry likely expired (TTL elapsed between calls) or cross-region inference routed to a Region without the entry. checkpoint-not-moved: The cached part was read, but the last cache point didn't move while 8 new blocks were added after it. Those are paid at full price on every call. Put the cache point at the end of the newest message. (messages[4].content[0]) $ cachecanary diff with-download.json without-download.json --model us.anthropic.claude-sonnet-5-5 ttl-changed: Checkpoint TTL changed (1h -> 5m); entries are keyed by TTL.
In plain words: the first pair is the 12-tool-call step from finding 2. The second is the role: "user" chat from finding 1: the cache works, but the newest messages are never added to it. The third is the same Sonnet 5.5 request with and without LiteLLM's price list download (finding 4).
Keep a few saved requests in your repo and run cachecanary lint on them in CI, and a LiteLLM upgrade that changes your cache points shows up in a pull request instead of on the bill. Setup is on the home page.
How I tested this
LiteLLM 1.104.0 in a clean Python environment. To capture requests, I pointed LiteLLM at a small local server standing in for Bedrock that saved every request body. For the live numbers I sent the same requests to Bedrock (Claude Sonnet 4.6, us-west-2, October 7 2026) with made-up prompts and a fresh random tag in each run, so old cache entries couldn't help. The capture and the live checks are scripts in the repo, so you can run them yourself: capture and live checks. I also ran LiteLLM's own proxy server against live Bedrock, called with the OpenAI client as shown above, and got the same numbers as the library. Token counts move by a few tokens from run to run because each run uses a new random tag. LiteLLM changes quickly, so check your own version with the callback above.
More on why caching breaks on Bedrock, with the bug reports behind each case: Five ways Claude prompt caching quietly breaks on Amazon Bedrock.