Claude Extended Thinking Budget Planner: budget_tokens vs max_tokens, Tool Use with Thinking Blocks Passed Back, Streaming Thresholds, Caching Effects, and a Cost Test Plan
Ppromptstudio·Oct 9, 2026
No rating
Plan how to turn on extended thinking in the Claude API for a real feature: choose a budget_tokens and max_tokens pair per task, keep thinking blocks intact across tool calls, know which parameters stop working, decide when to stream, and test cost and quality before rollout.
Act as an LLM platform engineer who adds extended thinking to Claude API features in production, sets thinking budgets per task from test data, and keeps tool use, caching, and billing behavior correct.
Inputs:
- The feature and its task types (for example contract clause analysis, multi step data question answering, code review), with current model and settings: [FeatureAndTasks]
- Current request shape: system prompt size, typical input tokens, max_tokens, temperature, tool definitions, and tool_choice: [CurrentRequest]
- Whether the feature uses tool calls in a loop and how the conversation history is stored: [ToolLoop]
- Latency and cost limits per request and per month: [Limits]
- Whether prompt caching is enabled and where the cache breakpoints are: [CachingSetup]
- Output format: [Format]
Generate:
1. A task by task decision on whether extended thinking is worth enabling for FeatureAndTasks, with the reason.
2. A budget table per task: starting budget_tokens, max_tokens, and why, keeping budget_tokens at the documented minimum or above and below max_tokens unless interleaved thinking applies [confirm in Anthropic docs].
3. Parameter conflicts in CurrentRequest that change when thinking is on: temperature and top_k changes, forced tool_choice, and assistant prefill, with the fix for each [confirm in Anthropic docs].
4. A tool loop plan from ToolLoop: pass the thinking blocks, including their signatures, back unchanged in the assistant turn that holds the tool_use, and whether interleaved thinking between tool calls is needed.
5. A streaming rule: when max_tokens is large enough that streaming should be used, and how the UI handles thinking deltas.
6. A caching note from CachingSetup: what changing the thinking budget does to cached prefixes and where to keep breakpoints.
7. A billing note: thinking tokens are billed as output tokens, and on newer models the returned thinking may be a summary while billing covers the full thinking [confirm per model].
8. A cost and quality test plan against Limits: a fixed eval set, two or three budgets per task, and the metrics to log (usage.output_tokens, latency, pass rate).
Constraints:
- Do not invent prices, model names, or limits; mark them [confirm in Anthropic docs].
- Never suggest editing or removing thinking blocks mid tool loop.
- Plain engineer tone. No em dashes.