Back to Discover

#claude prompt caching

1 prompt found

Claude API Prompt Caching Planner: Where to Put cache_control Breakpoints, TTL Choice, and How to Prove Cache Hits in Usage Fields
๐Ÿค– AI Tools

Claude API Prompt Caching Planner: Where to Put cache_control Breakpoints, TTL Choice, and How to Prove Cache Hits in Usage Fields

PpromptstudioยทOct 5, 2026
No rating

Plan prompt caching for an app on the Anthropic Messages API: reorder a request so stable content sits in the cached prefix, place up to four cache_control breakpoints across tools, system, and messages, choose the 5 minute or 1 hour TTL, and verify hits with cache_creation_input_tokens and cache_read_input_tokens.

Act as an LLM platform engineer who tunes Anthropic Messages API workloads for cost and latency with prompt caching. You know caching only pays when the prefix is byte for byte identical between calls, so most of your work is reordering the request, not adding flags. Inputs: - Current request structure: tools, system prompt parts, documents, conversation history, and the user turn, with rough token sizes for each: [RequestLayout] - Which parts change per call, per user, per session, or almost never: [Volatility] - Traffic shape (calls per minute, gaps between calls in one session, sessions per day): [Traffic] - Model in use and SDK or HTTP client: [ModelAndSdk] - Current usage block from a real response, if you have one: [UsageSample] - Anything that varies silently, such as timestamps, request ids, shuffled tool order, or user names in the system prompt: [HiddenVariation] - Output format: [Format] - Language: [Lang] Generate: 1. Prefix order. Caching follows the order tools, then system, then messages. Rewrite RequestLayout so the most stable content comes first, and move every item from HiddenVariation out of the prefix or make it deterministic. 2. Breakpoint plan. Place cache_control with type ephemeral on the last block of each stable segment, using at most four breakpoints, and say what each one caches. Check each cached prefix against the model's minimum cacheable length (it differs by model; confirm the current number in the docs for ModelAndSdk), because shorter prefixes are silently not cached. 3. TTL choice from Traffic: the default 5 minute TTL refreshes on every hit; the 1 hour TTL costs more to write and suits gaps longer than 5 minutes. Explain the trade in terms of write versus read multipliers and tell the user to confirm current pricing. 4. Request example: the JSON body for ModelAndSdk with the breakpoints in place, using placeholders for long content. 5. Verification: what cache_creation_input_tokens, cache_read_input_tokens, and input_tokens should look like on the first call, the second call in the window, and after expiry. Read UsageSample and diagnose it. 6. Invalidation list for this app: the changes that will break the cache (editing tools, changing tool_choice, adding or removing images, editing earlier messages) and a logging check to catch regressions. Constraints: - Do not quote prices or token minimums as fixed facts; mark them as values to confirm in current Anthropic docs. - Never move user specific private data into a shared cached prefix across users. - No em dashes.