🤖 AI Tools

Gemini API Context Caching Planner: Implicit vs Explicit Caching on Gemini 2.5 Models, google-genai caches.create with TTL, Minimum Cacheable Tokens, Stable Prefix Prompt Layout, Storage vs Input Cost Math, and Cache Hit Monitoring

Decide whether and how to cache a large repeated context in the Gemini API: compare implicit caching with explicit caches, lay out prompts so the stable part comes first, create and refresh a cache with the google-genai SDK and a TTL that matches your traffic, check the cost against storage charges using the rates you paste, and confirm hits in usage metadata.

0.0
0Reviews
P
October 7, 2026

Prompt

Act as a backend engineer who runs production apps on the Gemini API, has cut monthly token bills by caching large shared contexts, and has debugged the error you get when a request sends a system instruction alongside a cached context.

Inputs:
- The workload: requests per day, hours when traffic arrives, shared context size in tokens, and typical question size: [Workload]
- What repeats across requests (system instruction, manuals, codebase, a long video or PDF) and how often it changes: [SharedContext]
- Model and SDK in use: [ModelChoice]
- Rates copied from the current Gemini pricing page for input, cached input, and cache storage: [PricingNotes]
- Current prompt structure or code snippet: [CurrentCode]
- Output format: [Format]

Generate:
1. An implicit versus explicit decision: implicit caching is on by default for Gemini 2.5 models and gives a discount only when a request shares a prefix with a recent one, with no guarantee; explicit caching stores the context for a TTL you pay storage for, and gives a predictable discount. Pick one for this Workload and say why.
2. Prompt layout rules for either mode: put the large stable content first and the per request question last, keep the stable part byte identical, and send similar requests close together in time.
3. A minimum size check: explicit caches need a minimum number of input tokens that differs by model; tell the user to measure with client.models.count_tokens and compare against the current minimum for ModelChoice from the docs.
4. Python code with the google-genai SDK: client.caches.create with model, a CreateCachedContentConfig holding display_name, system_instruction, contents, and ttl; then client.models.generate_content with GenerateContentConfig(cached_content=cache.name). Put system_instruction and tools in the cache, not in the request, because the request cannot set them when it uses cached content.
5. TTL and refresh: choose a TTL from the traffic hours in Workload, extend it with client.caches.update when traffic continues, and delete it with client.caches.delete when the day ends.
6. Versioning: caches are not edited in place; create a new cache when SharedContext changes, and put the content version in display_name.
7. Cost math using only PricingNotes: daily cost without caching, with explicit caching (cached tokens at the cached rate plus storage tokens times hours), and the break even request count.
8. Monitoring: log usage_metadata.cached_content_token_count and prompt_token_count per request, alert when the cached share drops, and recreate the cache on a not found error after expiry.

Constraints:
- Never invent prices or token minimums; use PricingNotes and mark anything else "confirm in the current Gemini docs".
- The cache model must match the request model exactly. No em dashes.

Instructions

Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.

Generated Output

This image was generated using the prompt above.

Gemini API Context Caching Planner: Implicit vs Explicit Caching on Gemini 2.5 Models, google-genai caches.create with TTL, Minimum Cacheable Tokens, Stable Prefix Prompt Layout, Storage vs Input Cost Math, and Cache Hit Monitoring - Result

Examples

Example Input

Workload: support assistant, about 4,000 questions a day between 9 AM and 6 PM, each question about 200 tokens, answers about 400 tokens
SharedContext: 95,000 token product manual plus a 1,500 token system instruction; the manual changes about once a month
ModelChoice: gemini-2.5-flash with the google-genai Python SDK
PricingNotes: as pasted from the pricing page today: input 0.30 per 1M tokens, cached input 0.075 per 1M tokens, storage 1.00 per 1M tokens per hour
CurrentCode: the manual is pasted after the user's question in every request, with system_instruction set in each call

Example Output

1) DECISION: EXPLICIT CACHE
4,000 requests a day over 9 hours share a 96,500 token prefix. Implicit caching might catch many of these, but only if the prefix comes first and requests stay close in time, with no guarantee. An explicit cache gives a known discount every request. Use explicit, and fix the layout so implicit hits also happen during the switch.

2) LAYOUT FIX
Today the manual comes after the question, so no two requests share a prefix. Move the manual into the cache and send only the question per request.

3) SIZE CHECK
tokens = client.models.count_tokens(model="gemini-2.5-flash", contents=[manual_text])
96,500 tokens is far above any current minimum; confirm the exact minimum for gemini-2.5-flash in the current Gemini docs.

4) CODE
from google import genai
from google.genai import types
client = genai.Client()
cache = client.caches.create(
    model="gemini-2.5-flash",
    config=types.CreateCachedContentConfig(
        display_name="manual-v2026-10",
        system_instruction=SYSTEM_TEXT,
        contents=[manual_text],
        ttl="3600s",
    ),
)
resp = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=question,
    config=types.GenerateContentConfig(cached_content=cache.name),
)
Remove system_instruction from the per request config. Sending it with cached_content returns an error.

5) TTL AND REFRESH
- Create at 8:55 AM with a 1 hour TTL.
- A scheduler calls client.caches.update(name=cache.name, config=types.UpdateCachedContentConfig(ttl="3600s")) every 50 minutes until 6 PM.
- Call client.caches.delete(name=cache.name) at 6:05 PM so storage stops overnight.

6) VERSIONING
When the manual changes, create "manual-v2026-11", switch traffic, then delete the old cache.

7) COST MATH (prefix only; question and answer tokens cost the same either way)
- Without cache: 4,000 x 96,500 = 386M tokens x 0.30 per 1M = 115.80 per day
- With cache: 386M x 0.075 per 1M = 28.95, plus storage 96,500 tokens x 9.2 hours = 0.89M token hours x 1.00 = 0.89, plus one creation at full input price 0.03. Total about 29.87 per day
- Savings: about 85.93 per day at these pasted rates
- Break even: storage for one hour (0.0965) is covered after about 5 cached requests in that hour, since each saves 96,500 x 0.225 per 1M = 0.0217

8) MONITORING
log(resp.usage_metadata.prompt_token_count, resp.usage_metadata.cached_content_token_count)
- Alert if cached share falls under 90 percent for 10 minutes.
- On a not found error for the cache name, recreate the cache and retry once.

Reviews (0)

Please login to leave a review.
Loading reviews...