Claude Sonnet 5.5 and Opus 5.5: API migration and cache costs
Claude 5.5 migration needs a request and conversation-state review, not just a model-name replacement.
Published 2026-10-03 · Updated 2026-10-03 · KeepRouter Editorial · 7 minute read

Start Claude Sonnet 5.5 and Opus 5.5 with the native Messages endpoint on KeepRouter, then test a complete tool conversation. These models change thinking and sampling behavior. Removing an unsupported sampling parameter may fix an initial 400 response, but losing a thinking block can still break the next tool turn.
The exact Sonnet 5.5 and Opus 5.5 pages link to Anthropic's documentation and current customer prices. This guide reviews API behavior on October 3, 2026; it does not claim independent benchmark results.
What changes when moving from an older Claude?
Both models document a one-million-token context and up to 128,000 output tokens. Sonnet uses adaptive thinking by default; Opus keeps adaptive thinking enabled. Their default effort levels differ. Non-default temperature, top_p and top_k are not supported. For an old agent configuration, remove overrides first rather than searching for another temperature that produces the same behavior.
Inspect tool-choice settings too. A forced tool choice can conflict with adaptive thinking. Begin with automatic tool selection and a narrow schema, then test any stricter policy against the exact model documentation. Old computer-use tool versions also require a separate migration review. Messages compatibility is not a promise that every legacy beta tool is available.
Use the Anthropic SDK with a KeepRouter key
import os
from anthropic import Anthropic
client = Anthropic(base_url="https://keeprouter.com",
api_key=os.environ["KEEPROUTER_KEY"], max_retries=0)
r = client.messages.create(
model="claude-sonnet-5-5", max_tokens=1024,
output_config={"effort": "low"},
messages=[{"role": "user", "content": "Describe one safe database rollback rule."}])
print(r.content)
print(r.usage)Use the origin as the SDK base URL; the SDK appends /v1/messages. Raw HTTP requests use a KeepRouter key in x-api-key, an anthropic-version header and JSON content. The quickstart explains account and key setup. OpenAI-compatible chat is also listed for these IDs, but native Messages is the clearer starting point when preserving Anthropic thinking and tool blocks matters.
Preserve the assistant message between tool turns
Suppose your agent asks a read-only get_invoice tool for an invoice amount. Keep the entire returned assistant content, including the thinking block and its signature, before adding the tool-result user message. Match the tool-use ID exactly. Do not rebuild an assistant turn as a plain text summary or strip blocks to save a few prompt tokens.
A practical failure test deliberately returns a missing invoice. The model should report that absence instead of inventing a balance. Another test returns a decimal amount and a currency; the final answer must preserve both. These examples are suggested acceptance tasks, not claims about a customer deployment.
Understand the fixed customer input rate
Anthropic distinguishes ordinary input, cache creation at different lifetimes and cache reads. KeepRouter's published 5.5 input rate is one fixed customer ceiling covering ordinary input and the supported one-hour cache-write rate. The same price also applies to ordinary and five-minute-write input. It is therefore higher than the maker's ordinary-input base price; no separate cache-write surcharge is added.
Cache hits are charged at the displayed cached-input price. Treat the full prompt as a partition: ordinary input plus cache creation plus cache reads. Reconcile the returned usage fields with the request ledger instead of adding a cache-read subtotal to an input total that already contains it. The cache audit guide explains why apparent savings can disappear after a changing prefix or retries.
For repeated repository work, compare an uncached first request, a subsequent identical prefix, and a request whose prefix has changed. Record actual cache usage in each case. Never present a projected cache ratio as a measured saving.
Choose Sonnet or Opus with an acceptance rule
Take a fixed set of repository fixes or document decisions. Define acceptance before seeing either answer: tests pass, cited evidence is present, the required tool was used and no unauthorized action occurred. Use the same tool schemas and context. Start at low effort for the baseline, then increase it only when a failure has a plausible reasoning cause.
Compare total billed spend per accepted result, including failed attempts and manual repairs, and report latency from your own application. A larger model name does not prove higher reliability for your task. Conversely, a low token rate does not prove lower cost if the agent needs extra turns.
Keep the previous model available during a small migration cohort. Roll back when accepted-result rate falls or tool history fails; an authentication failure should be fixed at the key level. Create a key, run the text example and then the invoice tool loop before switching a production coding agent.
Primary documentation checked for this guide:Anthropic · Sonnet 5.5 · Anthropic · Opus 5.5.
Frequently asked questions
Can old temperature settings be copied to Claude 5.5?
Non-default temperature, top_p and top_k are unsupported. Remove those overrides and review the exact model guide.
Why preserve thinking blocks?
They are part of the model conversation state. Rebuilding tool history as plain text can lose information required by the next turn.
Does cache support always reduce the bill?
No. Confirm returned cache usage and reconcile all attempts at the published customer rates.
Sources reviewed
Article last reviewed 2026-10-03