Gemini API on KeepRouter: choose a model, connect and check costs
Start with a model suited to your task, use one KeepRouter client and inspect the returned usage. This guide maps exact model IDs to evaluation tasks and shows what we tested.
Published 2026-09-29 · Updated 2026-10-03 · KeepRouter Editorial · 6 minute read

To call Gemini on KeepRouter, use a KeepRouter API key, the base URL https://keeprouter.com/v1, and an exact model ID from the catalog. Start with Flash-Lite for a classification or extraction evaluation, Flash for a reasoning or coding evaluation, and Pro Preview when the task warrants a separate complex-reasoning test. Check the current price on each model page before running a batch.
Is the Gemini endpoint the same as other chat models?
Yes: the customer-facing request is POST https://keeprouter.com/v1/chat/completions, authenticated with your KeepRouter key. Keep the same OpenAI client and change model to the exact public Gemini ID. The gateway manages Google authentication and the upstream publisher mapping, so your application does not pass a Google Cloud project, region or credential.
This compatibility applies to Chat Completions. The current Google Cloud chat route does not expose Responses, Anthropic Messages, Google-native generateContent or Gemini Live WebSockets. A model reachable through another route can have a different endpoint set; check its model page. For streaming, JSON schemas and tools, use the model-specific parameters and tested examples below.
Choose a model for a concrete task
The nine chat models below returned text in our September 29, 2026 KeepRouter API checks. The suggested tasks are evaluation starting points, not a measured ranking. Use your own examples and acceptance criteria when comparing models.
| Exact model ID and current price | Start by evaluating |
|---|---|
| gemini-3.8-flash | Reasoning and coding with low reasoning effort |
| gemini-3.7-flash | The same tasks as 3.8, comparing quality and token consumption |
| gemini-3.6-flash | An existing evaluation that needs a pinned Flash version |
| gemini-3.5-flash | Extraction, coding and tool-use fixtures from an existing application |
| gemini-3.5-flash-lite | Frequent classification and schema-constrained extraction |
| gemini-3.1-flash-lite | Existing Flash-Lite tasks with adjustable thinking |
| gemini-3.1-pro-preview | Complex reasoning, with a review of preview lifecycle risks |
| gemini-3.1-pro-preview-customtools | Custom-tool conversations that need a complete tool round trip |
| gemini-3-flash-preview | An existing Gemini 3 Flash preview integration |
For a support-ticket extractor, define success as a valid schema plus correct values for ticket category, priority and requested action. For coding, run the generated change against the same tests. A reply arriving successfully is only the first check. Measure retries and human corrections too; they affect the cost of a usable result.
Make a first request with the OpenAI SDK
Follow the quickstart to create a KeepRouter key and store it in a server-side environment variable. Use the exact ID below, without adding a provider prefix. This example uses Chat Completions and disables automatic retries so the first test is easy to reconcile.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://keeprouter.com/v1",
api_key=os.environ["KEEPROUTER_API_KEY"],
timeout=60.0,
max_retries=0,
)
reply = client.chat.completions.create(
model="gemini-3.8-flash",
messages=[{"role": "user", "content": "Write a short title for a login-error ticket."}],
reasoning_effort="low",
max_tokens=1024,
)
print(reply.choices[0].message.content)
print(reply.choices[0].finish_reason)
print(reply.usage)The output cap must leave room for both thinking and the final answer. A very small cap can be consumed before visible text appears. Gemini 3.8 Flash accepts low, medium and high reasoning effort; minimal is not supported. Keep an explicit model ID when changing parameters, because one generation's settings do not establish another's behavior. See the API reference for the request contract and the migration checker for client configuration checks.
Know what has been tested
These checks ran through KeepRouter on September 29, 2026. They are small integration fixtures, not latency or throughput benchmarks.
| Check | Exact scope | Observed result |
|---|---|---|
| Text response | All nine IDs in the table above | Each returned a successful answer |
| Fully consumed stream | Gemini 3.8 Flash | Text, final usage and stream completion received |
| JSON-schema response | Gemini 3.5 Flash-Lite | Returned the requested {"ok":true} object |
| Tool conversation | Gemini 3.1 Pro Preview Custom Tools | Tool call followed by a successful continuation using its result |
For streaming, set stream=True and stream_options={"include_usage": True} and consume the stream to completion. Text chunks alone do not establish final usage. For JSON, validate the returned object in your application. For tools, retain the complete assistant message, call IDs and any non-text fields before appending the matching tool result. Validate arguments before executing a tool.
Media inputs, Gemini Live audio and Google-native tools require their own compatibility checks. The text and tool fixtures above do not establish support for those separate operations. If your application depends on them, check the selected model's documented endpoint before switching traffic.
Calculate cost from returned usage
Use the current KeepRouter input, output and cached-input rates displayed on the exact model page. The cost calculator estimates a planned workload; the console usage view records actual requests.
For a token-priced request with reported cache hits, the calculation is:
cost = (input tokens - cached input tokens) / 1,000,000 × input rate
+ cached input tokens / 1,000,000 × cached-input rate
+ completion tokens / 1,000,000 × output rateUse a cached-input rate only where the model publishes one and the response reports cache hits. Completion usage already includes reported reasoning tokens; do not add the reasoning detail a second time. In our Custom Tools continuation fixture, the response reported 40 input tokens, 42 completion tokens and 82 total tokens. The 35 reasoning tokens were part of those 42 completion tokens, not an extra chargeable quantity.
Compare cost per accepted result as well as cost per request. If two models differ in retry frequency or need different prompt lengths, their per-million-token prices alone do not answer which is cheaper for your task. Run the same small batch, count accepted results and divide total recorded spend by that count. Keep failed and incomplete requests visible in the comparison.
Select the next test from the feature you need
If the application extracts fields, test schema validity and factual accuracy on several document shapes. If it uses tools, test a complete continuation. If it streams, test clean completion and cancellation. The bounded results above help choose a starting point; they do not rank the models. Google’s Gemini 3.8 Flash model documentation describes the model, while the linked KeepRouter model page defines the customer route and price. Use the cost calculator with your expected input and output before a larger paid trial.
Move an application in a small, reversible step
Save a baseline from your current model: inputs, acceptance checks and typical usage. Run those examples against one selected Gemini ID, then compare correctness and total consumption. Include a completed stream or tool conversation if your application uses it. Keep preview models behind a configuration flag so a model change does not require rewriting the application.
After the small batch passes, increase traffic gradually and review errors, latency and spend in the console. Use request observability to connect a failed user action to its request record. Keep the previous model configuration available until the new workload has passed your own acceptance checks.
Frequently asked questions
Do Gemini requests need a Google key, and do all OpenAI endpoints work?
Use a KeepRouter key with /v1/chat/completions and the public Gemini model ID. Google authentication stays on the gateway. The current Google Cloud chat route does not expose /v1/responses, /v1/messages, generateContent or Live WebSockets.
How do I call Gemini with the OpenAI SDK on KeepRouter?
Set base_url to https://keeprouter.com/v1, use a KeepRouter key and an exact catalog model ID. These examples use chat.completions.create.
Should I add reasoning tokens to completion tokens?
No. Reported reasoning is already included in completion usage. Use the completion count once when applying the output rate.
Which Gemini features were tested on KeepRouter?
On September 29, 2026: text on all nine listed IDs, streaming on 3.8 Flash, JSON on 3.5 Flash-Lite and a tool round trip on 3.1 Pro Preview Custom Tools. This is a bounded integration check.
Where can I check the current Gemini API price?
Each KeepRouter model page shows its current customer rates. Compare the same workload and include retries when estimating cost per accepted result.
Sources reviewed
Article last reviewed 2026-10-03
- [1] KeepRouter API reference
- [2] KeepRouter models and current prices
- [3] Gemini 3.8 Flash model card
- [4] Gemini 3.1 Pro and custom tools