How to Use the Claude API Prompt Caching Planner to Cut Repeated Input Costs
Plan prompt caching for an app on the Anthropic Messages API: reorder a request so stable content sits in the cached prefix, place up to four cache_control breakpoints across tools, system, and messages, choose the 5 minute or 1 hour TTL, and verify hits with cache_creation_input_tokens and cache_read_input_tokens.

Many Claude API apps send the same large chunk of input on every call: a long system prompt, a policy manual, a set of tool definitions. Prompt caching lets the API reuse that prefix instead of processing it fresh each time, which lowers cost and latency for the cached part. The catch is that the cached prefix must be exactly the same between calls. One timestamp or a user name near the top of the system prompt, or tools that come out in a different order, and the cache never hits. The Claude API Prompt Caching Planner: Where to Put cache_control Breakpoints, TTL Choice, and How to Prove Cache Hits in Usage Fields prompt focuses on that reordering work first, then places the breakpoints.
What the prompt produces
- A new prefix order that puts stable content first, following the tools, system, then messages order, and moves anything that varies out of the cached part.
- A breakpoint plan with up to four cache_control markers and what each one caches.
- A TTL choice between the default five minute cache and the longer one hour option, based on your traffic.
- An example request body for your SDK with the breakpoints in place.
- A verification guide for reading cache_creation_input_tokens, cache_read_input_tokens, and input_tokens.
- An invalidation list of changes that will break the cache in your app, plus a logging check.
How to fill the inputs
RequestLayout describes your request piece by piece with rough token sizes: tools, system prompt parts, documents, history, and the new user turn. Volatility says how often each part changes. Traffic gives calls per minute and the gaps between turns in one session, which decides whether the cache stays warm.
ModelAndSdk names the model and the client library. UsageSample is a real usage block from a response; it is the fastest way to see whether caching works today. HiddenVariation is where you list anything that changes quietly, like a date, a request id, or a user name in the system prompt.
Reading the example output
The example is a support bot with a large policy manual, fourteen tools, and ticket history. Its usage sample shows zero cache activity. Lessons worth copying:
- Find the hidden variation first. The date and customer name sat at the start of the system prompt, and the tools were sometimes built in a different order. Either one stops the prefix from being reused.
- Split the system prompt. The static role and manual become one block with a breakpoint; the date and name move to a small block after it.
- Sort the tools every time. Deterministic order makes the tool definitions cacheable.
- Cache shared and per user content separately. The tools and manual are shared across all users; the ticket history is cached only for that ticket's own calls, so no private data lands in a shared prefix.
- Choose TTL per breakpoint. Constant traffic keeps the shared prefix warm on the default TTL, while ticket turns that arrive minutes apart may justify the longer one.
- Prove it in the usage fields. The first call shows cache creation, the next calls show cache reads, and input tokens drop to the uncached tail.
Tips for better results
Start from a real usage block. It tells you right away whether anything is cached.
Confirm numbers in the docs. Minimum cacheable lengths, TTL options, and pricing multipliers vary by model and can change. The prompt marks these as values to check rather than stating them as fixed.
Log a hash of the prefix. If cache reads stay at zero after a change, a changing hash shows which part still varies.
Add a test for regressions. A simple test that fails when the cached block contains a date or a name stops someone from breaking the cache in a later template edit.
Mistakes to avoid
- Adding cache_control without reordering the request.
- Putting user specific data in a prefix shared across users.
- Assuming a short prefix is cached when it is below the model minimum.
- Editing earlier messages or switching tool_choice and wondering why hits stop.
Who it is for
This prompt fits backend engineers running Claude in production, teams building support bots or document assistants with large system prompts, and anyone whose API bill is mostly repeated input. It is also a good review tool before a launch, to catch prefix problems early.
Related PromptDig links
- Prompt: Claude API Prompt Caching Planner: Where to Put cache_control Breakpoints, TTL Choice, and How to Prove Cache Hits in Usage Fields
- More AI tools prompts: Browse more prompts
- Have an LLM engineering prompt to share? Share a prompt