Find out when your Claude prompt cache stops working on Bedrock.
Bedrock can reuse the start of your prompt, like your instructions and tool list, and charge about 10% of the normal price for it. When that quietly stops, every request still works and you pay full price again. Usually the bill is the first sign. CacheCanary shows how often your requests use the cache and what the misses cost you, tells you why a request missed, and catches the problem in CI before it ships.
Free and open source (Apache 2.0). The tool runs on your machine and sends nothing anywhere.
Framework guides: LiteLLM · Strands Agents · LangChain
What a broken cache costs
The example uses Anthropic's list price for Claude Sonnet 4.6 ($3 per million input tokens). Bedrock prices can vary by region.
How you run it
- Install it with
pip install cachecanary. It runs on your laptop or in GitHub Actions. No account, no signup, no server. - Check a request: save it as a JSON file and run one command. No AWS access needed.
- See your real hit rate and what misses cost: copy your Bedrock log files from S3 with one AWS command, then point CacheCanary at the folder.
- Safe to try:
lint,diffandlogsonly read files. Onlyprobecalls Bedrock, and only with your own AWS credentials.
How prompt caching works, in 30 seconds
- You add a cache point to your request, usually right after your instructions and tools. Everything before it is the part Bedrock can reuse.
- That part has to be exactly the same on every request. Change one character, like today's date, and Bedrock can't reuse it.
- It also has to be long enough: at least 1,024 tokens on Sonnet 4.6 and 4,096 on Haiku 4.5. Shorter prompts are never cached, and nothing tells you.
- A cached prompt lasts 5 minutes (or 1 hour if you ask for it), and every request that uses it starts the clock again.
What CacheCanary does
1. Shows whether you're losing money right now
Turn on Bedrock's invocation logging (it's off by default), copy the log files to your machine and point CacheCanary at the folder. For each model it shows how often requests used the cache, what the input cost, and how much of that the misses cost you.
$ aws s3 sync s3://YOUR-BUCKET/AWSLogs/YOUR-ACCOUNT-ID/BedrockModelInvocationLogs/ bedrock-logs/ $ cachecanary logs bedrock-logs/ --min-hit 0.8 us.anthropic.claude-sonnet-4-6: hit 50% over 8 calls (read 11496, write 11496, uncached 80, unparsed 0) <-- below threshold input cost $0.05 at list price; about $0.03 of it lost to cache misses (target: 80% hit rate) List prices: Amazon Bedrock on-demand, Oct 2026 (global. IDs at list, others +10%). Use --price for your own rate.
Here only half the calls used the cache, below the 80% target, so the command fails and shows what the misses cost on this small sample. Run it as a scheduled CI job and that failure becomes your alert. Prices are Bedrock's published list prices; pass --price if you pay a different rate.
2. Tells you why a request missed the cache
Save a request your app sends to Bedrock as a JSON file. With boto3's Converse API that's one line:
# kwargs: what you pass to client.converse(**kwargs)
json.dump(kwargs, open("request.json", "w"))
Then check it, or compare it with an earlier one:
$ cachecanary lint request.json [warn] dynamic-in-prefix: Found ISO date text inside the cached prefix. If it changes per request, every call misses. (system[0]) $ cachecanary diff yesterday.json today.json system-changed: The system prompt changed inside the cached prefix (look for dates, IDs or per-user text). (system[0]) $ cachecanary lint request.json --model us.anthropic.claude-haiku-4-5-20251001-v1:0 [error] prefix-too-short: Prefix up to this checkpoint is ~1892 tokens; claude-haiku-4-5 needs at least 4096. The request succeeds but nothing is cached. (system[0])
In plain words: the first request has a date in its system prompt, so it changes every day. The second shows the system prompt changed between yesterday and today. The third prompt is too short for Haiku 4.5, so it's never cached, even though the same prompt caches fine on Sonnet 4.6. system[0] points to the first block of the system prompt.
3. Catches it before it ships
Add it to GitHub Actions. Problems show up as comments on the pull request, next to the file that caused them.
- uses: Haarris/cachecanary@v0
with:
command: lint
args: tests/fixtures/agent_request.json --model us.anthropic.claude-sonnet-4-6
If your CI has AWS access, probe goes further: it sends the request to Bedrock twice and fails if the second one wasn't served from the cache.
Why caching breaks
Often nobody changed the prompt on purpose. These are the usual causes:
- A library upgrade moved the system prompt or dropped the cache point. This happened with LiteLLM in July 2026: cache hits fell from about 90% to 25-45% and spend went up 2-3x for six days.
- You switched to a new model ID or an inference profile (an AWS name that points to a model), and your library doesn't recognise it, so it stops asking for caching.
- A small edit put today's date, a request ID or the user's name into the system prompt.
- Your tools are listed in a different order on each request.
- Thinking or effort settings changed between two calls. On Bedrock that throws away the whole cache, system prompt included.
- An agent ran a dozen tools at once. Bedrock only searches back about 20 pieces of the conversation (messages, tool calls and tool results) for the cached part, so it can lose track of it.
I wrote these up in more detail, with the bug reports behind each one and the lookback limit I measured: Five ways Claude prompt caching quietly breaks on Amazon Bedrock.
Anthropic added cache diagnostics to its own API, so you can see why a request missed. They don't cover Bedrock. That's the gap CacheCanary fills.
Guides for your framework
I captured exactly what each framework sends to Bedrock, checked it with CacheCanary and tested it on live Bedrock. Each guide shows where the cache is lost and a setup that keeps it.
- Prompt caching with LiteLLM on Amazon Bedrock:
role: "user"caching your oldest messages, wide tool calls, and a 1-hour cache lifetime that is quietly dropped. - Prompt caching with Strands Agents on Amazon Bedrock: caching that is off by default, wide tool calls, and long chats that lose the cache on every turn.
- Prompt caching with LangChain on Amazon Bedrock: ChatBedrock losing the whole cache after 10 tool calls, context editing that rewrites the conversation on every call, and your own cache points turning off LangChain's extra one.
Built on how Bedrock actually behaves
The checks follow AWS's documented caching rules. The core behaviour was verified against live Amazon Bedrock, including streaming responses, both cache lifetimes (5 minutes and 1 hour), and each model's minimum prompt size. The log reader is tested against real Bedrock invocation logs.
Coming next
A hosted dashboard: cache hit rate and wasted spend per app, an alert when the rate drops, and the reason for each miss. It runs inside your own AWS account, so prompts never leave it.
If you'd use it, email hello@cachecanary.com. I read every message.