How to Use the Gemini Context Caching Prompt to Cut Token Costs on Large Repeated Contexts
Decide whether and how to cache a large repeated context in the Gemini API: compare implicit caching with explicit caches, lay out prompts so the stable part comes first, create and refresh a cache with the google-genai SDK and a TTL that matches your traffic, check the cost against storage charges using the rates you paste, and confirm hits in usage metadata.

If your app sends the same long manual, codebase, or transcript to the Gemini API with every request, you are paying full input price for the same tokens again and again. Context caching can change that, but only if you pick the right mode, lay out the prompt correctly, set a sensible TTL, and confirm that requests are actually hitting the cache. The Gemini API Context Caching Planner: Implicit vs Explicit Caching on Gemini 2.5 Models, google-genai caches.create with TTL, Minimum Cacheable Tokens, Stable Prefix Prompt Layout, Storage vs Input Cost Math, and Cache Hit Monitoring prompt walks through all of that for your specific workload and writes the google-genai code to go with it.
What the prompt produces
- An implicit versus explicit decision for your workload, with the reason.
- Prompt layout rules so the stable content always comes first.
- A minimum size check using token counting, with a reminder to compare against the current minimum for your model.
- Python code with the google-genai SDK to create a cache and use it in requests.
- A TTL and refresh plan that matches your traffic hours, including when to delete the cache.
- A versioning approach for when the shared content changes.
- Cost math using the rates you paste from the pricing page.
- Monitoring based on the usage metadata in each response.
How to fill the inputs
Workload describes requests per day, the hours they arrive, the size of the shared context in tokens, and the typical question size. Traffic hours drive the TTL plan.
SharedContext names what repeats and how often it changes, such as a product manual that updates monthly.
ModelChoice gives the model and SDK. The cache model has to match the request model exactly, so be precise.
PricingNotes should be copied from the current pricing page on the day you run the prompt. The prompt does not invent prices.
CurrentCode shows how you build requests today. This is often where the biggest win hides.
Reading the example output
The example is a support assistant answering about 4,000 questions a day against a 95,000 token manual on gemini-2.5-flash.
- The layout is the first fix. The current code puts the manual after the question, so no two requests share a prefix and implicit caching can never help.
- Explicit caching is chosen because the traffic is steady across business hours and a predictable discount is worth paying storage for.
- The system instruction moves into the cache. Sending a system instruction alongside cached content causes an error, so the request only carries the question.
- The TTL follows the workday. The cache is created just before traffic starts, extended through the day, and deleted after hours so storage stops.
- The math uses only the pasted rates. It compares the daily cost with and without caching and shows how few requests per hour cover the storage cost.
Tips for better results
- Count tokens with the SDK instead of estimating from page counts.
- Put a version in the cache display name so you always know which content a request used.
- Log cached token counts from day one, not after the bill arrives.
- Rerun the cost math whenever pricing or traffic changes.
Mistakes to avoid
- Do not put the variable question before the stable content.
- Do not edit shared content and expect the cache to update. Create a new cache.
- Do not leave caches alive overnight if nobody uses them.
- Do not assume a cache hit. Check the usage metadata.
Who it is for
Backend engineers and AI product teams running retrieval or support assistants on Gemini, developers building tools over long documents or codebases, and technical founders watching their API bill.
Related PromptDig links
Open the Gemini API Context Caching Planner: Implicit vs Explicit Caching on Gemini 2.5 Models, google-genai caches.create with TTL, Minimum Cacheable Tokens, Stable Prefix Prompt Layout, Storage vs Input Cost Math, and Cache Hit Monitoring prompt and describe your workload to get a caching plan and code. For more AI tool and developer prompts, Browse more prompts. If you have an LLM engineering prompt that saved you time or money, Share a prompt.