# GPT API pricing and SDK migration: keep the request contract visible

> Moving a GPT integration changes a billing destination and can change its API contract. Start with one existing text flow, preserve its endpoint family, verify the exact model ID, and reconcile usage before moving tools or stateful conversations.

_Published 2026-09-23 · Updated 2026-09-29 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 5 minute read_

![Field-by-field OpenAI-compatible API migration checklist from request contract to rollback](https://keeprouter.com/editorial/blog/openai-compatible-api-migration-checklist.png)

_Keep endpoint, model, parser and usage checks together when migrating a client._

A GPT API migration has two separate budgets: the cost of generated work and the cost of changing the integration. A lower published input rate is useful only if the new route accepts the fields your application sends and returns the objects it reads. Start with one existing feature, such as summarizing a support ticket, and write down its current request and response contract.

This guide assumes the application already uses an OpenAI SDK. KeepRouter exposes compatible routes, but the chosen model still needs an active route for the operation. Use the [model catalog](/models), including an entry such as [GPT-4o](/models/gpt-4o), to identify a candidate. The example is a configuration pattern, not evidence of a completed live test or a recommendation to select a particular model generation.

## Separate endpoint migration from model replacement

There are three independent changes: the host and key, the model ID, and the API family. Changing all three at once makes a failure difficult to diagnose. First reproduce a basic text feature through the intended endpoint; then add any model change; finally evaluate a different API family if the feature needs it.

| Existing dependency | What to inventory | Evidence needed after the change |
|---|---|---|
| Chat Completions | messages, choices, finish_reason | Correct text and terminal state |
| Responses | input items, output items, event types | Parser accepts each required item |
| Tool loop | Function schema, IDs and tool results | One complete call and continuation |
| Conversation state | Stored history or response identifiers | Next turn retains the required context |
| Usage | Input, cached input, output, other charges | Cost reconciles with the account record |

OpenAI's migration documentation explains the different object and streaming models of Responses and Chat Completions. A method named `responses.create` in your SDK does not prove that a gateway model has the corresponding configured route. Keep the [Responses comparison](/blog/responses-api-vs-chat-completions) beside the inventory and inspect the [public API contract](/api/docs) for the target operation.

## Make one bounded Python request

Set `KEEPROUTER_MODEL` to an exact current chat model from the catalog and `KEEPROUTER_KEY` to a key allowed to call it. Pin the OpenAI Python package version in your application's lockfile. This snippet uses Chat Completions and avoids sampling and reasoning options.

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://keeprouter.com/v1",
    api_key=os.environ["KEEPROUTER_KEY"],
    timeout=30.0,
    max_retries=0,
)
result = client.chat.completions.create(
    model=os.environ["KEEPROUTER_MODEL"],
    messages=[{"role": "user", "content": "Summarize: the test ticket is resolved."}],
    max_completion_tokens=128,
)
print(result.id, result.model)
print(result.choices[0].finish_reason)
print(result.usage.model_dump() if result.usage else "usage missing")
```

Confirm that the selected route accepts `max_completion_tokens` before adopting this example. If its contract requires a different output-limit field, change that field deliberately. Disabling SDK retries during the first diagnostic run helps identify one attempt; it is not a production reliability recommendation. A later retry policy should account for deadlines and the possibility that an earlier attempt already performed billable work.

Notice that the key is read from the environment. A Python string containing a shell variable name is only a literal string. If authentication fails, check presence and scope without printing the credential. The [SDK setup guide](/use-cases/openai-sdk) includes the base URL spelling for other clients.

## Build the GPT cost worksheet from usage

For token-priced text routes, calculate uncached input plus cached input plus output using the service's published rates. For OpenAI-style prompt totals, subtract reported cached tokens from total input once. Keep any separately billed tool, image, audio, batch or service-tier charges outside this simple formula unless the selected service explicitly includes them.

Suppose a feature runs 40,000 requests per month, each averaging 1,500 input and 250 output tokens. With no cache assumption, that is 60 million input and 10 million output tokens. The model portion is therefore `60 * input_price_per_million + 10 * output_price_per_million`. This is a planning example, not observed traffic. Run a second scenario with your longer requests and measure the actual distribution before relying on an average.

Use the [API cost calculator](/tools/api-cost-calculator) to change the token assumptions. An advertised context window is a maximum capability, not the amount charged on every request. Conversely, passing a previous response identifier does not make historical context inherently free. Confirm the usage fields the actual operation reports.

## Expand the canary in a useful order

After a basic request succeeds, add streaming and ensure the application handles the final event and usage. Next test one tool invocation with a harmless deterministic result. Then test the second turn of the conversation. Finally, exercise a rejected model ID and an expired or out-of-scope key in a controlled environment so the application displays useful errors.

For each case, record the request shape, SDK version, route, status, completion condition and observed charge. Do not compare exact prose when the task only requires preserving a ticket's meaning. Use a small checklist of required facts and unacceptable inventions instead. If the new route changes tool behavior, keep that feature on the previous route while adapting its parser.

## Recheck cache accounting when changing model generations

Do not assume every GPT generation uses the same cache-write treatment. [OpenAI’s current caching guide](https://developers.openai.com/api/docs/guides/prompt-caching) documents model-specific behavior. For a KeepRouter call, use the customer rate and quantities defined for the exact model and route; native-provider cache options should not be assumed to pass through. Test a cache miss and a supported cache read, and account for any separately billed write before projecting savings.

## Keep a rollback that includes state and billing

Store the previous host, credential reference, model ID and parser choice as one deployable configuration. Rolling back only the URL can leave the application sending the new model namespace or reading the wrong event shape. Keep outstanding stateful conversations on a compatible path until they finish or migrate their history explicitly.

Review the first paid sample for cost per accepted summary, including retries and rejected answers. A successful HTTP response, a correct answer and a settled account charge answer different questions. When those records agree, expand one feature at a time and revisit the worksheet when model prices or user behavior change.

## Frequently asked questions

### Does changing base_url migrate every OpenAI feature?

No. Verify the selected model, endpoint, tools, stream events and state handling for each feature.

### Do I need to switch to Responses during a host migration?

Only if the feature requires it. Keeping the current API family first makes a host migration easier to diagnose.

### Can the calculator predict my final GPT invoice?

It estimates the selected published units. Real usage, retries and separately billed features determine the final charge.

### Why use an environment lookup for the API key?

The Python SDK needs the secret value. A literal shell-variable string is not expanded by Python.

## Sources reviewed

_Article last reviewed 2026-09-29_

1. [OpenAI Responses migration guide](https://developers.openai.com/api/docs/guides/migrate-to-responses)
2. [OpenAI Python SDK](https://github.com/openai/openai-python)
3. [OpenAI prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching)
4. [KeepRouter OpenAPI](https://keeprouter.com/api/openapi.json)

## Related guides

- [gpt 4o](https://keeprouter.com/models/gpt-4o.md)
- [openai sdk](https://keeprouter.com/use-cases/openai-sdk.md)
- [Responses API vs Chat Completions: choose by contract, not novelty](https://keeprouter.com/blog/responses-api-vs-chat-completions.md)
- [api cost calculator](https://keeprouter.com/tools/api-cost-calculator)
- [Gemini with the OpenAI SDK: compare the native, compatible and gateway paths](https://keeprouter.com/blog/gemini-openai-compatible-api-differences.md)
- [OpenRouter to KeepRouter migration: map URLs, model IDs and routing fields](https://keeprouter.com/blog/openrouter-to-keeprouter-migration.md)

## Check your request before migrating

Check model IDs and request fields in your browser. No API key or inference call is needed.

[Check a sample configuration](https://keeprouter.com/tools/api-migration-checker)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
