🤖 AI Tools
Claude API Prompt Caching Planner: Where to Put cache_control Breakpoints, TTL Choice, and How to Prove Cache Hits in Usage Fields
Plan prompt caching for an app on the Anthropic Messages API: reorder a request so stable content sits in the cached prefix, place up to four cache_control breakpoints across tools, system, and messages, choose the 5 minute or 1 hour TTL, and verify hits with cache_creation_input_tokens and cache_read_input_tokens.
0Reviews
Prompt
Act as an LLM platform engineer who tunes Anthropic Messages API workloads for cost and latency with prompt caching. You know caching only pays when the prefix is byte for byte identical between calls, so most of your work is reordering the request, not adding flags. Inputs: - Current request structure: tools, system prompt parts, documents, conversation history, and the user turn, with rough token sizes for each: [RequestLayout] - Which parts change per call, per user, per session, or almost never: [Volatility] - Traffic shape (calls per minute, gaps between calls in one session, sessions per day): [Traffic] - Model in use and SDK or HTTP client: [ModelAndSdk] - Current usage block from a real response, if you have one: [UsageSample] - Anything that varies silently, such as timestamps, request ids, shuffled tool order, or user names in the system prompt: [HiddenVariation] - Output format: [Format] - Language: [Lang] Generate: 1. Prefix order. Caching follows the order tools, then system, then messages. Rewrite RequestLayout so the most stable content comes first, and move every item from HiddenVariation out of the prefix or make it deterministic. 2. Breakpoint plan. Place cache_control with type ephemeral on the last block of each stable segment, using at most four breakpoints, and say what each one caches. Check each cached prefix against the model's minimum cacheable length (it differs by model; confirm the current number in the docs for ModelAndSdk), because shorter prefixes are silently not cached. 3. TTL choice from Traffic: the default 5 minute TTL refreshes on every hit; the 1 hour TTL costs more to write and suits gaps longer than 5 minutes. Explain the trade in terms of write versus read multipliers and tell the user to confirm current pricing. 4. Request example: the JSON body for ModelAndSdk with the breakpoints in place, using placeholders for long content. 5. Verification: what cache_creation_input_tokens, cache_read_input_tokens, and input_tokens should look like on the first call, the second call in the window, and after expiry. Read UsageSample and diagnose it. 6. Invalidation list for this app: the changes that will break the cache (editing tools, changing tool_choice, adding or removing images, editing earlier messages) and a logging check to catch regressions. Constraints: - Do not quote prices or token minimums as fixed facts; mark them as values to confirm in current Anthropic docs. - Never move user specific private data into a shared cached prefix across users. - No em dashes.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
RequestLayout: system = 'You are SupportBot for Kestrel Bikes. Today is {date}. Customer: {first_name}.' (40 tokens) + policy manual (9,000 tokens) + 14 tool definitions (3,500 tokens); messages = ticket history (0 to 6,000 tokens) + new customer message (100 to 400 tokens)
Volatility: manual and tools change monthly, ticket history grows per turn, customer message changes every call
Traffic: about 30 calls per minute across all users in business hours; one ticket gets 3 to 8 turns, often 2 to 15 minutes apart
ModelAndSdk: Claude Sonnet via the Python anthropic SDK
UsageSample: input_tokens 12,980, cache_creation_input_tokens 0, cache_read_input_tokens 0 on the 50th call of the day
HiddenVariation: the date and first name in line one of the system prompt; tools are built from a dict and sometimes come out in a different order
Format: sections 1 to 6
Lang: EnglishExample Output
1. Prefix order
Your sample shows zero cache activity because no breakpoints are set yet, and two things would defeat them anyway: tools come first and sometimes change order, which changes the start of the prefix, and {date} and {first_name} open the system prompt, so nothing after them could be reused. New order:
- tools: 14 definitions, sorted by name every time (fixes the dict ordering)
- system block 1: the static SupportBot role line without date or name, plus the policy manual
- system block 2: Today is {date}. Customer: {first_name}. (small, uncached)
- messages: ticket history, then the new customer message
2. Breakpoints
- BP1 on the last tool: caches about 3,500 tokens shared by every user.
- BP2 on system block 1: caches tools plus manual, about 12,500 tokens, shared by every user. This is the big win.
- BP3 on the last message of ticket history: caches the ticket so turn 4 reuses turns 1 to 3. Per ticket, so it only pays when turns arrive inside the TTL.
- One breakpoint unused. All prefixes are well above the minimum length for Sonnet; confirm the current minimum in the docs.
Private data check: BP1 and BP2 contain no customer data; BP3 caches one customer's own ticket and is only reused by that ticket's calls.
3. TTL
- BP1 and BP2: 30 calls per minute keeps the shared prefix warm all day, so the default 5 minute TTL is enough.
- BP3: turns are 2 to 15 minutes apart, so many arrive after 5 minutes. Use the 1 hour TTL here only if the extra write cost is less than the reads it saves; with 3 to 8 turns per ticket it usually is, but confirm against current write and read multipliers. Longer TTL breakpoints must come before shorter ones in the prefix, so if you mix, verify the ordering rule in the docs.
4. Request body (Python SDK)
client.messages.create(model="<sonnet model id>", max_tokens=1024,
tools=[*sorted_tools[:-1], {**sorted_tools[-1], "cache_control": {"type": "ephemeral"}}],
system=[{"type": "text", "text": ROLE_AND_MANUAL, "cache_control": {"type": "ephemeral"}},
{"type": "text", "text": f"Today is {date}. Customer: {first_name}."}],
messages=[*history_with_cache_on_last_block, {"role": "user", "content": new_msg}])
5. Verification
- First call after deploy: cache_creation_input_tokens about 12,500 plus the ticket, cache_read_input_tokens 0.
- Next call from any user within 5 minutes: cache_read_input_tokens about 12,500; input_tokens drops to the uncached tail (system block 2 plus new message).
- Next turn on the same ticket within the TTL: cache reads also include that ticket's earlier history.
- If reads stay 0 after the fix, log a hash of the serialized tools and system block 1 per call; a changing hash means something still varies.
6. What breaks the cache here
- Monthly manual or tool edits (expected; one write per deploy).
- Changing tool_choice, adding an image to a ticket, or editing an earlier message in history.
- Any template change that puts the date back into block 1. Add a test that fails if ROLE_AND_MANUAL contains a date or the customer name.