← Learn
Playbook · 10 steps · checked 4 October 2026 · 0 adopted

API

At most ten steps, in order of how much they help. Each step is from Anthropic's own pages, linked. Re-checked against those pages every week; anything that can't be confirmed is marked unverified rather than removed.

  1. 01

    Pick a model with your own tests, pin its ID, and plan for retirement

    Anthropic suggests starting with Claude Opus 5.5 (claude-opus-5-5) for most workloads, or starting efficiency-first with Claude Haiku 4.5 and moving up only if your tests show a gap. Every model ID is a pinned snapshot, including dateless ones such as claude-sonnet-5-5; Anthropic ships updates under a new ID. Older aliases such as claude-sonnet-4-5 do move, so use the full ID. Keep the ID in config, not scattered through code. Models are deprecated with at least 60 days' notice, and requests to a retired model fail. For example, claude-sonnet-4-5-20250929 retires on 30 November 2026, with claude-sonnet-5-5 as the replacement. To find old models in use, export the CSV on the Console Usage page, which breaks usage down by API key and model. Re-run your tests before you switch. (models overview, choosing a model, model IDs and versioning, model deprecations, migration guides)

  2. 02

    Build an evaluation set before you tune anything

    Anthropic calls a good evaluation set "the most important step" in deciding whether to change models. Write specific, measurable success criteria, often across several dimensions such as accuracy, tone, latency and cost. Collect test cases that look like real traffic, plus edge cases: irrelevant, overly long or ambiguous input. Grade automatically where you can, with code checks or a model as grader. Use a different model to grade than the one that produced the output. Many cases with rough automatic grading beat a few graded by hand. Run the set on every change of model, effort level or prompt. (define success and build evals, choosing a model)

  3. 03

    Cache the parts of the request that do not change

    Put stable content first, in the order tools, then system, then messages, and anything that changes per request (timestamps, the user's message) after it. Add a top-level cache_control: {"type": "ephemeral"} for automatic caching in conversations, or place cache_control on the last block that is identical across requests. The cache lasts 5 minutes by default; "ttl": "1h" gives an hour. Prompts below a minimum length are not cached and no error is returned, so check cache_read_input_tokens and cache_creation_input_tokens in usage. Changing tools, thinking settings or top-level output_config.effort breaks the cache from that point. For most models, cache reads also do not count towards your input-tokens-per-minute rate limit. Caches are kept separate per workspace. (prompt caching, rate limits, workspaces)

  4. 04

    Get machine-readable output with structured outputs and strict tools

    When your code parses Claude's reply, set output_config.format with a JSON schema rather than parsing free text. When Claude calls your functions, add strict: true to each tool so its inputs always match the schema. Both are generally available with no beta header. The SDKs can build the schema for you: client.messages.parse() with Pydantic in Python, or Zod in TypeScript. Know the limits: no recursive schemas, additionalProperties must be false, no numeric or string length constraints, and at most 20 strict tools per request. A refusal or a max_tokens stop can still return output that does not match, so check stop_reason before you parse. For tools, either run the loop yourself (tool_use in, tool_result back) or let the SDK's Tool Runner do it. (structured outputs, tool use overview, stop reasons)

  5. 05

    Set effort on purpose and give max_tokens room

    Effort (output_config.effort: low, medium, high, xhigh, max) is the main control for cost, speed and depth on current models. It affects all output, including thinking and tool calls. Defaults differ: medium on Claude Opus 5.5, high on most others, so set it explicitly and choose it with an effort sweep on your evaluation set rather than copying an old setting. Use low for simple, high-volume or latency-sensitive calls. max_tokens is a hard limit on thinking plus reply, so set it large at higher effort. Thinking is adaptive and always on for some models; sending thinking: {"type": "disabled"} to them returns a 400 error. From Claude Opus 4.7 onwards, setting temperature, top_p or top_k to a non-default value also returns a 400 error. (effort, models overview, errors, model deprecations)

  6. 06

    Stream anything long

    Set "stream": true, or use the SDK's messages.stream() helper, so users see text as it arrives. For long or large-max_tokens requests, streaming is not optional: the SDKs refuse non-streaming requests expected to run past 10 minutes, and idle connections can be dropped. If you do not need the text as it arrives, call get_final_message() (TypeScript: finalMessage()) to get the complete message over a stream. Errors can arrive mid-stream after a 200 response, so handle error events too. To resume an interrupted stream on Claude 4.6 and later, send the partial text back in a user message and ask Claude to continue. (streaming, errors: long requests)

  7. 07

    Handle errors, rate limits and stop reasons properly

    The SDKs retry connection errors, 429s and 5xx errors twice by default with exponential backoff and honour retry-after; set max_retries to suit your app. Catch the SDK's typed exceptions rather than matching message text. A 529 means the API is overloaded. A 429 with error_code enforced_spend_limit_reached and no retry-after is the monthly spend cap, and retrying will not help. Limits use a token bucket, so short bursts can trip them: ramp traffic up gradually and watch the anthropic-ratelimit-* response headers. Log the request-id header for support. Check stop_reason on every response: max_tokens means the reply was cut off, refusal means Claude declined, and tool_use means your code must run a tool. (errors, rate limits, stop reasons)

  8. 08

    Send work that can wait to the Message Batches API

    Batches suit bulk jobs such as evaluations, classification and back-fills. They cost 50% less than standard calls and have their own rate limits. Most batches finish within an hour; any request not done in 24 hours expires and is not billed. A batch holds up to 100,000 requests or 256 MB. Give each request a unique custom_id, because results may not come back in order. Download results within 29 days. With shared context across a batch, use the 1-hour cache lifetime, since batches often take longer than 5 minutes. (batch processing, rate limits)

  9. 09

    Measure tokens and cap spend

    The token counting endpoint (messages.count_tokens) takes the same input as a message and is free, with its own rate limit; the count is an estimate. Claude Opus 4.7 and later use a newer tokenizer that gives roughly 30% more tokens for the same text, so recount when you migrate rather than reusing old figures. Log the usage object from each response. Total input is cache_read_input_tokens + cache_creation_input_tokens + input_tokens. Set your own spend limit under Settings > Billing, and lower spend and rate limits per workspace so one project cannot use up another's share. The Console Usage page shows your cache hit rate. (token counting, rate limits, workspaces)

  10. 10

    Keep API keys on the server, scoped and short-lived

    Never ship a key in browser or mobile code. Read it from the ANTHROPIC_API_KEY environment variable or a secrets manager, rotate it, and disable or delete any key you think has leaked. Set an expiry when you create a key. Use a service account key for shared or production workloads, not a personal key, which stops working when its owner leaves; older workspace keys are now legacy. For production on a cloud platform or in CI, Workload Identity Federation swaps static keys for short-lived tokens. For iOS and macOS apps that call Claude directly, App Attest issues short-lived tokens to genuine installs. Use separate workspaces for development, staging and production, each with its own keys and limits. (authentication, workspaces)

Unverified

What Anthropic's own pages disagree on, or could not be confirmed.

  • Claude Haiku 4.5's retirement is listed as "not sooner than October 15, 2026", eleven days after this check. No deprecation notice was on the deprecations page on 2026-10-04. If you start efficiency-first on Haiku 4.5, watch that page.
  • The tool use overview says you can require a tool call by setting tool_choice. The errors page says Claude Opus 5.5, Claude Sonnet 5.5 and Claude Fable 5.1 reject tool_choice of type any or tool with a 400 error, and suggests auto with strict tools or structured outputs instead. Step 4 follows the errors page.
  • The structured outputs page names output_config.format as the parameter and calls output_format deprecated, but its Python messages.parse() example passes output_format=. Check your SDK version's docs for the exact helper argument.
  • Changing effort partway through a conversation without breaking the cache (a per-message output_config) is in beta and only on some models, so it is left out of Step 5.
  • Fast mode (speed: "fast") is a research preview with premium pricing and separate rate limits; it was not read in full, so it is not a step.
  • The multi-model "advisor" and "orchestrator" patterns on the cost and intelligence page were not read, so they are not a step. (choosing a model links to them.)