🤖 AI Tools

Claude Extended Thinking Budget Planner: budget_tokens vs max_tokens, Tool Use with Thinking Blocks Passed Back, Streaming Thresholds, Caching Effects, and a Cost Test Plan

Plan how to turn on extended thinking in the Claude API for a real feature: choose a budget_tokens and max_tokens pair per task, keep thinking blocks intact across tool calls, know which parameters stop working, decide when to stream, and test cost and quality before rollout.

0.0
0Reviews
P
October 9, 2026

Prompt

Act as an LLM platform engineer who adds extended thinking to Claude API features in production, sets thinking budgets per task from test data, and keeps tool use, caching, and billing behavior correct.

Inputs:
- The feature and its task types (for example contract clause analysis, multi step data question answering, code review), with current model and settings: [FeatureAndTasks]
- Current request shape: system prompt size, typical input tokens, max_tokens, temperature, tool definitions, and tool_choice: [CurrentRequest]
- Whether the feature uses tool calls in a loop and how the conversation history is stored: [ToolLoop]
- Latency and cost limits per request and per month: [Limits]
- Whether prompt caching is enabled and where the cache breakpoints are: [CachingSetup]
- Output format: [Format]

Generate:
1. A task by task decision on whether extended thinking is worth enabling for FeatureAndTasks, with the reason.
2. A budget table per task: starting budget_tokens, max_tokens, and why, keeping budget_tokens at the documented minimum or above and below max_tokens unless interleaved thinking applies [confirm in Anthropic docs].
3. Parameter conflicts in CurrentRequest that change when thinking is on: temperature and top_k changes, forced tool_choice, and assistant prefill, with the fix for each [confirm in Anthropic docs].
4. A tool loop plan from ToolLoop: pass the thinking blocks, including their signatures, back unchanged in the assistant turn that holds the tool_use, and whether interleaved thinking between tool calls is needed.
5. A streaming rule: when max_tokens is large enough that streaming should be used, and how the UI handles thinking deltas.
6. A caching note from CachingSetup: what changing the thinking budget does to cached prefixes and where to keep breakpoints.
7. A billing note: thinking tokens are billed as output tokens, and on newer models the returned thinking may be a summary while billing covers the full thinking [confirm per model].
8. A cost and quality test plan against Limits: a fixed eval set, two or three budgets per task, and the metrics to log (usage.output_tokens, latency, pass rate).

Constraints:
- Do not invent prices, model names, or limits; mark them [confirm in Anthropic docs].
- Never suggest editing or removing thinking blocks mid tool loop.
- Plain engineer tone. No em dashes.

Instructions

Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.

Generated Output

This image was generated using the prompt above.

Claude Extended Thinking Budget Planner: budget_tokens vs max_tokens, Tool Use with Thinking Blocks Passed Back, Streaming Thresholds, Caching Effects, and a Cost Test Plan - Result

Examples

Example Input

FeatureAndTasks: in app assistant for a property management SaaS; tasks: (1) short FAQ answers, (2) lease clause risk review from pasted lease text, (3) rent roll questions that call a SQL tool several times; currently a Claude Sonnet model with no thinking
CurrentRequest: system prompt about 3,500 tokens; inputs 1,000 to 30,000 tokens; max_tokens 4,096; temperature 0.2; tools run_sql and get_unit; tool_choice {"type": "any"} on the rent roll task
ToolLoop: yes, up to 6 tool calls per question; history stored in our DB as JSON message arrays, we currently drop any block that is not text before saving
Limits: p95 latency under 40 seconds for lease review, under 20 seconds for rent roll; budget about 0.08 USD per lease review
CachingSetup: cache_control breakpoint after the system prompt and tool definitions
Format: decisions, budget table, conflicts, tool loop, streaming, caching, billing, test plan

Example Output

1) DECISIONS
- FAQ answers: leave thinking off. Short answers do not gain much and latency matters.
- Lease clause review: turn thinking on. Multi clause reasoning over long text benefits most.
- Rent roll questions: turn thinking on with interleaved thinking, so the model can reason between SQL results [confirm beta header in Anthropic docs].

2) BUDGET TABLE
| Task | budget_tokens | max_tokens | Why |
| Lease review | 6,000 | 12,000 | room for reasoning plus a 2 to 3 page answer |
| Rent roll | 3,000 per turn | 8,000 | several short reasoning steps between tools |
budget_tokens stays at or above the documented minimum of 1,024 and below max_tokens [confirm in Anthropic docs].

3) PARAMETER CONFLICTS
- temperature 0.2: thinking does not support custom temperature; remove it for thinking requests [confirm].
- tool_choice {"type": "any"} on rent roll: forced tool use is not supported with thinking; switch to {"type": "auto"} and tell the model in the system prompt to query before answering.
- No assistant prefill is used today, keep it that way for thinking requests.

4) TOOL LOOP
Your DB layer drops non text blocks. That will break the loop: the assistant message with tool_use must be sent back with its thinking block and signature unchanged. Change storage to save the full content array. Do not edit, trim, or reorder thinking blocks. Older thinking blocks from earlier completed turns can be left in; the API handles them [confirm].

5) STREAMING
Stream the lease review and rent roll tasks. Large max_tokens values require or strongly favor streaming to avoid request timeouts [confirm threshold in Anthropic docs]. Show a "Reviewing the lease" status while thinking deltas arrive; do not show raw thinking to tenants.

6) CACHING
Keep the breakpoint after system prompt and tools. Changing budget_tokens between requests can invalidate cached message content, so fix one budget per task and do not tune it per request [confirm].

7) BILLING
Thinking tokens bill as output tokens. On newer Claude models the returned thinking may be summarized while billing covers the full thinking, so read usage.output_tokens instead of counting visible text [confirm per model].

8) TEST PLAN
- Eval set: 40 lease excerpts with lawyer marked risks; 40 rent roll questions with known answers.
- Budgets to compare: lease 3,000, 6,000, 10,000; rent roll 1,500, 3,000, 5,000.
- Log usage.input_tokens, usage.output_tokens, cache read tokens, p95 latency, and pass rate.
- Pick the smallest budget within 2 points of the best pass rate that meets the 40 second and 20 second p95 targets and the 0.08 USD lease cost, priced from the current price page [confirm].

Reviews (0)

Please login to leave a review.
Loading reviews...