GPT-6.1 Sol API: Responses, tools and a reliable migration
Use Responses for GPT-6.1 Sol function calling. Choose a model after testing a complete tool loop and the cost of an accepted result.
Published 2026-10-03 · Updated 2026-10-03 · KeepRouter Editorial · 7 minute read

Use POST /v1/responses for GPT-6.1 Sol function calling on KeepRouter. Its Chat Completions interface does not support tools, and its reasoning cannot be turned off. A working text response therefore does not prove that your existing agent integration will work. Start with the exact gpt-6.1-sol model page, a bounded output budget and one harmless tool.
This guide was checked against OpenAI's exact model pages on October 3, 2026. Model limits are documentation facts. The evaluation below is a reproducible method, not a published benchmark or a throughput guarantee.
Choose the endpoint before the model
| Public model ID | Start here | Tool migration issue |
|---|---|---|
| gpt-6.1-sol | Responses | Do not send tools through Chat Completions |
| gpt-6-sol | Responses | Chat tools require non-thinking mode |
| gpt-6-luna | Responses | Chat tools require non-thinking mode |
All three have separate public price IDs. Sol and Luna are not names for the same product at different temperatures. For an existing client, check whether it constructs input or messages, where it places reasoning effort, and how it returns tool results. The migration checker helps inspect those request differences.
Make one small Responses request
Store the key on the server. This Python example uses the OpenAI SDK, disables automatic retries for reconciliation, and chooses low effort rather than an unsupported no-thinking mode on 6.1 Sol.
import os
from openai import OpenAI
client = OpenAI(base_url="https://keeprouter.com/v1",
api_key=os.environ["KEEPROUTER_KEY"], max_retries=0)
r = client.responses.create(
model="gpt-6.1-sol",
input="Explain what a database transaction does in one sentence.",
reasoning={"effort": "low"}, max_output_tokens=1024, store=False)
print(r.output_text)
print(r.usage)The output limit covers thinking as well as the answer. If the result is incomplete, inspect the completion status before raising the budget. Do not keep retrying an unchanged request that cannot fit its answer.
Verify a tool round trip, not just a tool name
Use a local lookup_order function with a single string argument and a fixed test response. No purchase, email or account change is needed. A passing test has four parts: the model emits a function call, the arguments match your schema, the application returns the result with the matching call ID, and the model produces an answer grounded in that result.
When using store=False, keep the returned output items in your application's conversation history. Add a function_call_output item for the result and send the continuing input through Responses. Do not discard reasoning items merely because your chat UI does not display them. Provider-managed conversation state is a separate capability; do not assume a gateway persists it for you.
Log a request ID, model ID, status, finish state, total input, cached input and output usage. Keep API keys and customer payloads out of general telemetry. A visible final answer with the correct order value is more useful evidence than HTTP 200 alone.
Calculate the price of an accepted change
Use the model page's current KeepRouter customer rates. The catalog uses fixed ceilings covering the supported long-context and cache-write tiers. These differ from OpenAI's short-context base list rates, and ordinary input uses the same published ceiling. There is no second cache-write surcharge. A cache hit uses the separately displayed cached-input rate.
For token-metered work, calculate (input - cached input) × input rate + cached input × cache rate + output × output rate. Output includes reported reasoning once; do not add the reasoning breakdown again. Sum all billable attempts, then divide by the number of changes that pass your tests. A cheap request that needs three repairs can cost more than a single accepted result.
The cost calculator lets you change prompt size, output size and cache assumptions. Keep an uncached scenario beside the expected cache scenario. A model having cache support does not mean your particular request had a cache hit.
Compare Sol and Luna on your own workload
Select twenty representative tasks: simple extraction, a small code fix, a repository change requiring two tools and a long-document question with a cited answer. This is a suggested evaluation size, not a claim that we ran twenty tasks. Hold tool schemas, test data and acceptance rules constant. Record accepted results, retries, billed usage and latency from your application.
If Luna passes the simpler tasks, route those tasks there explicitly and reserve Sol for tasks where it demonstrates a useful improvement. Do not automatically retry an authentication error with a more expensive model; changing the model does not fix an invalid key. For production, define a spend ceiling and a fallback that preserves the required endpoint and conversation history.
Create a KeepRouter key, run the bounded request, then complete the harmless tool loop before moving customer traffic. The Responses versus Chat Completions guide explains the request families in more detail.
Primary documentation checked for this guide:OpenAI · GPT-6.1 Sol · OpenAI · GPT-6 Sol · OpenAI · GPT-6 Luna.
Frequently asked questions
Can GPT-6.1 Sol use tools through Chat Completions?
No. Use Responses for function calling. A successful text reply does not verify tools.
Does the displayed output usage include reasoning?
Reported reasoning is included in completion usage once. Do not add its breakdown again.
Are KeepRouter rates identical to maker base rates?
No. These catalog IDs use published fixed ceilings for supported cache-write and long-context tiers. Read the live model price.
Sources reviewed
Article last reviewed 2026-10-03