GPT API pricing and SDK migration: keep the request contract visible

Moving a GPT integration changes a billing destination and can change its API contract. Start with one existing text flow, preserve its endpoint family, verify the exact model ID, and reconcile usage before moving tools or stateful conversations.

Published 2026-09-23 · Updated 2026-09-29 · KeepRouter Editorial · 5 minute read

Field-by-field OpenAI-compatible API migration checklist from request contract to rollback
Keep endpoint, model, parser and usage checks together when migrating a client.

A GPT API migration has two separate budgets: the cost of generated work and the cost of changing the integration. A lower published input rate is useful only if the new route accepts the fields your application sends and returns the objects it reads. Start with one existing feature, such as summarizing a support ticket, and write down its current request and response contract.

This guide assumes the application already uses an OpenAI SDK. KeepRouter exposes compatible routes, but the chosen model still needs an active route for the operation. Use the model catalog, including an entry such as GPT-4o, to identify a candidate. The example is a configuration pattern, not evidence of a completed live test or a recommendation to select a particular model generation.

Separate endpoint migration from model replacement

There are three independent changes: the host and key, the model ID, and the API family. Changing all three at once makes a failure difficult to diagnose. First reproduce a basic text feature through the intended endpoint; then add any model change; finally evaluate a different API family if the feature needs it.

Existing dependencyWhat to inventoryEvidence needed after the change
Chat Completionsmessages, choices, finish_reasonCorrect text and terminal state
Responsesinput items, output items, event typesParser accepts each required item
Tool loopFunction schema, IDs and tool resultsOne complete call and continuation
Conversation stateStored history or response identifiersNext turn retains the required context
UsageInput, cached input, output, other chargesCost reconciles with the account record

OpenAI's migration documentation explains the different object and streaming models of Responses and Chat Completions. A method named responses.create in your SDK does not prove that a gateway model has the corresponding configured route. Keep the Responses comparison beside the inventory and inspect the public API contract for the target operation.

Make one bounded Python request

Set KEEPROUTER_MODEL to an exact current chat model from the catalog and KEEPROUTER_KEY to a key allowed to call it. Pin the OpenAI Python package version in your application's lockfile. This snippet uses Chat Completions and avoids sampling and reasoning options.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://keeprouter.com/v1",
    api_key=os.environ["KEEPROUTER_KEY"],
    timeout=30.0,
    max_retries=0,
)
result = client.chat.completions.create(
    model=os.environ["KEEPROUTER_MODEL"],
    messages=[{"role": "user", "content": "Summarize: the test ticket is resolved."}],
    max_completion_tokens=128,
)
print(result.id, result.model)
print(result.choices[0].finish_reason)
print(result.usage.model_dump() if result.usage else "usage missing")

Confirm that the selected route accepts max_completion_tokens before adopting this example. If its contract requires a different output-limit field, change that field deliberately. Disabling SDK retries during the first diagnostic run helps identify one attempt; it is not a production reliability recommendation. A later retry policy should account for deadlines and the possibility that an earlier attempt already performed billable work.

Notice that the key is read from the environment. A Python string containing a shell variable name is only a literal string. If authentication fails, check presence and scope without printing the credential. The SDK setup guide includes the base URL spelling for other clients.

Build the GPT cost worksheet from usage

For token-priced text routes, calculate uncached input plus cached input plus output using the service's published rates. For OpenAI-style prompt totals, subtract reported cached tokens from total input once. Keep any separately billed tool, image, audio, batch or service-tier charges outside this simple formula unless the selected service explicitly includes them.

Suppose a feature runs 40,000 requests per month, each averaging 1,500 input and 250 output tokens. With no cache assumption, that is 60 million input and 10 million output tokens. The model portion is therefore 60 * input_price_per_million + 10 * output_price_per_million. This is a planning example, not observed traffic. Run a second scenario with your longer requests and measure the actual distribution before relying on an average.

Use the API cost calculator to change the token assumptions. An advertised context window is a maximum capability, not the amount charged on every request. Conversely, passing a previous response identifier does not make historical context inherently free. Confirm the usage fields the actual operation reports.

Expand the canary in a useful order

After a basic request succeeds, add streaming and ensure the application handles the final event and usage. Next test one tool invocation with a harmless deterministic result. Then test the second turn of the conversation. Finally, exercise a rejected model ID and an expired or out-of-scope key in a controlled environment so the application displays useful errors.

For each case, record the request shape, SDK version, route, status, completion condition and observed charge. Do not compare exact prose when the task only requires preserving a ticket's meaning. Use a small checklist of required facts and unacceptable inventions instead. If the new route changes tool behavior, keep that feature on the previous route while adapting its parser.

Recheck cache accounting when changing model generations

Do not assume every GPT generation uses the same cache-write treatment. OpenAI’s current caching guide documents model-specific behavior. For a KeepRouter call, use the customer rate and quantities defined for the exact model and route; native-provider cache options should not be assumed to pass through. Test a cache miss and a supported cache read, and account for any separately billed write before projecting savings.

Keep a rollback that includes state and billing

Store the previous host, credential reference, model ID and parser choice as one deployable configuration. Rolling back only the URL can leave the application sending the new model namespace or reading the wrong event shape. Keep outstanding stateful conversations on a compatible path until they finish or migrate their history explicitly.

Review the first paid sample for cost per accepted summary, including retries and rejected answers. A successful HTTP response, a correct answer and a settled account charge answer different questions. When those records agree, expand one feature at a time and revisit the worksheet when model prices or user behavior change.

Frequently asked questions

Does changing base_url migrate every OpenAI feature?

No. Verify the selected model, endpoint, tools, stream events and state handling for each feature.

Do I need to switch to Responses during a host migration?

Only if the feature requires it. Keeping the current API family first makes a host migration easier to diagnose.

Can the calculator predict my final GPT invoice?

It estimates the selected published units. Real usage, retries and separately billed features determine the final charge.

Why use an environment lookup for the API key?

The Python SDK needs the secret value. A literal shell-variable string is not expanded by Python.

Sources reviewed

Article last reviewed 2026-09-29

  1. [1] OpenAI Responses migration guide
  2. [2] OpenAI Python SDK
  3. [3] OpenAI prompt caching
  4. [4] KeepRouter OpenAPI

Related guides

← All posts · Models & pricing · Get an API key