CacheCanary

Find out when your Claude prompt cache stops working on Bedrock.

Bedrock can reuse the start of your prompt, like your instructions and tool list, and charge about 10% of the normal price for it. When that quietly stops, every request still works and you pay full price again. Usually the bill is the first sign. CacheCanary shows how often your requests use the cache and what the misses cost you, tells you why a request missed, and catches the problem in CI before it ships.

pip install cachecanary

Free and open source (Apache 2.0). The tool runs on your machine and sends nothing anywhere.

Framework guides: LiteLLM · Strands Agents · LangChain

What a broken cache costs

90% offis what cached prompt text costs on Bedrock. When the cache breaks, that text goes back to full price.
$2,700 a monthextra for one agent with a 10,000-token prompt and 100,000 requests a month.
2-3xdaily spend for LiteLLM users for six days in July 2026, after one library update.

The example uses Anthropic's list price for Claude Sonnet 4.6 ($3 per million input tokens). Bedrock prices can vary by region.

How you run it

How prompt caching works, in 30 seconds

What CacheCanary does

1. Shows whether you're losing money right now

Turn on Bedrock's invocation logging (it's off by default), copy the log files to your machine and point CacheCanary at the folder. For each model it shows how often requests used the cache, what the input cost, and how much of that the misses cost you.

$ aws s3 sync s3://YOUR-BUCKET/AWSLogs/YOUR-ACCOUNT-ID/BedrockModelInvocationLogs/ bedrock-logs/
$ cachecanary logs bedrock-logs/ --min-hit 0.8
us.anthropic.claude-sonnet-4-6: hit 50% over 8 calls (read 11496, write 11496, uncached 80, unparsed 0)  <-- below threshold
  input cost $0.05 at list price; about $0.03 of it lost to cache misses (target: 80% hit rate)
List prices: Amazon Bedrock on-demand, Oct 2026 (global. IDs at list, others +10%). Use --price for your own rate.

Here only half the calls used the cache, below the 80% target, so the command fails and shows what the misses cost on this small sample. Run it as a scheduled CI job and that failure becomes your alert. Prices are Bedrock's published list prices; pass --price if you pay a different rate.

2. Tells you why a request missed the cache

Save a request your app sends to Bedrock as a JSON file. With boto3's Converse API that's one line:

# kwargs: what you pass to client.converse(**kwargs)
json.dump(kwargs, open("request.json", "w"))

Then check it, or compare it with an earlier one:

$ cachecanary lint request.json
[warn] dynamic-in-prefix: Found ISO date text inside the cached prefix. If it changes per request, every call misses. (system[0])

$ cachecanary diff yesterday.json today.json
system-changed: The system prompt changed inside the cached prefix (look for dates, IDs or per-user text). (system[0])

$ cachecanary lint request.json --model us.anthropic.claude-haiku-4-5-20251001-v1:0
[error] prefix-too-short: Prefix up to this checkpoint is ~1892 tokens; claude-haiku-4-5 needs at least 4096. The request succeeds but nothing is cached. (system[0])

In plain words: the first request has a date in its system prompt, so it changes every day. The second shows the system prompt changed between yesterday and today. The third prompt is too short for Haiku 4.5, so it's never cached, even though the same prompt caches fine on Sonnet 4.6. system[0] points to the first block of the system prompt.

3. Catches it before it ships

Add it to GitHub Actions. Problems show up as comments on the pull request, next to the file that caused them.

- uses: Haarris/cachecanary@v0
  with:
    command: lint
    args: tests/fixtures/agent_request.json --model us.anthropic.claude-sonnet-4-6

If your CI has AWS access, probe goes further: it sends the request to Bedrock twice and fails if the second one wasn't served from the cache.

Why caching breaks

Often nobody changed the prompt on purpose. These are the usual causes:

I wrote these up in more detail, with the bug reports behind each one and the lookback limit I measured: Five ways Claude prompt caching quietly breaks on Amazon Bedrock.

Anthropic added cache diagnostics to its own API, so you can see why a request missed. They don't cover Bedrock. That's the gap CacheCanary fills.

Guides for your framework

I captured exactly what each framework sends to Bedrock, checked it with CacheCanary and tested it on live Bedrock. Each guide shows where the cache is lost and a setup that keeps it.

Built on how Bedrock actually behaves

The checks follow AWS's documented caching rules. The core behaviour was verified against live Amazon Bedrock, including streaming responses, both cache lifetimes (5 minutes and 1 hour), and each model's minimum prompt size. The log reader is tested against real Bedrock invocation logs.

21Claude models with their cache rules built in, including Sonnet 4.6, Haiku 4.5 and Opus 5.5
Both APIsConverse and InvokeModel, streaming or not
No AWS neededto check a saved request. lint and diff run offline in about 30 ms.

Coming next

A hosted dashboard: cache hit rate and wasted spend per app, an alert when the rate drops, and the reason for each miss. It runs inside your own AWS account, so prompts never leave it.

If you'd use it, email hello@cachecanary.com. I read every message.