# Gemini API on KeepRouter: choose a model, connect and check costs

> Start with a model suited to your task, use one KeepRouter client and inspect the returned usage. This guide maps exact model IDs to evaluation tasks and shows what we tested.

_Published 2026-09-29 · Updated 2026-10-03 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 6 minute read_

![Field-by-field OpenAI-compatible API migration checklist from request contract to rollback](https://keeprouter.com/editorial/blog/openai-compatible-api-migration-checklist.png)

_Compatibility is approved operation by operation, including streams, tools, errors, usage, and rollback._

**To call Gemini on KeepRouter, use a KeepRouter API key, the base URL https://keeprouter.com/v1, and an exact model ID from the catalog.** Start with Flash-Lite for a classification or extraction evaluation, Flash for a reasoning or coding evaluation, and Pro Preview when the task warrants a separate complex-reasoning test. Check the current price on each model page before running a batch.

## Is the Gemini endpoint the same as other chat models?

Yes: the customer-facing request is `POST https://keeprouter.com/v1/chat/completions`, authenticated with your KeepRouter key. Keep the same OpenAI client and change `model` to the exact public Gemini ID. The gateway manages Google authentication and the upstream publisher mapping, so your application does not pass a Google Cloud project, region or credential.

This compatibility applies to Chat Completions. The current Google Cloud chat route does not expose Responses, Anthropic Messages, Google-native generateContent or Gemini Live WebSockets. A model reachable through another route can have a different endpoint set; check its model page. For streaming, JSON schemas and tools, use the model-specific parameters and tested examples below.

## Choose a model for a concrete task

The nine chat models below returned text in our September 29, 2026 KeepRouter API checks. The suggested tasks are evaluation starting points, not a measured ranking. Use your own examples and acceptance criteria when comparing models.

| Exact model ID and current price | Start by evaluating |
|---|---|
| [gemini-3.8-flash](/models/gemini-3.8-flash) | Reasoning and coding with low reasoning effort |
| [gemini-3.7-flash](/models/gemini-3.7-flash) | The same tasks as 3.8, comparing quality and token consumption |
| [gemini-3.6-flash](/models/gemini-3.6-flash) | An existing evaluation that needs a pinned Flash version |
| [gemini-3.5-flash](/models/gemini-3.5-flash) | Extraction, coding and tool-use fixtures from an existing application |
| [gemini-3.5-flash-lite](/models/gemini-3.5-flash-lite) | Frequent classification and schema-constrained extraction |
| [gemini-3.1-flash-lite](/models/gemini-3.1-flash-lite) | Existing Flash-Lite tasks with adjustable thinking |
| [gemini-3.1-pro-preview](/models/gemini-3.1-pro-preview) | Complex reasoning, with a review of preview lifecycle risks |
| [gemini-3.1-pro-preview-customtools](/models/gemini-3.1-pro-preview-customtools) | Custom-tool conversations that need a complete tool round trip |
| [gemini-3-flash-preview](/models/gemini-3-flash-preview) | An existing Gemini 3 Flash preview integration |

For a support-ticket extractor, define success as a valid schema plus correct values for ticket category, priority and requested action. For coding, run the generated change against the same tests. A reply arriving successfully is only the first check. Measure retries and human corrections too; they affect the cost of a usable result.

## Make a first request with the OpenAI SDK

Follow the [quickstart](/docs/quickstart) to create a KeepRouter key and store it in a server-side environment variable. Use the exact ID below, without adding a provider prefix. This example uses Chat Completions and disables automatic retries so the first test is easy to reconcile.

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://keeprouter.com/v1",
    api_key=os.environ["KEEPROUTER_API_KEY"],
    timeout=60.0,
    max_retries=0,
)
reply = client.chat.completions.create(
    model="gemini-3.8-flash",
    messages=[{"role": "user", "content": "Write a short title for a login-error ticket."}],
    reasoning_effort="low",
    max_tokens=1024,
)
print(reply.choices[0].message.content)
print(reply.choices[0].finish_reason)
print(reply.usage)
```

The output cap must leave room for both thinking and the final answer. A very small cap can be consumed before visible text appears. Gemini 3.8 Flash accepts low, medium and high reasoning effort; minimal is not supported. Keep an explicit model ID when changing parameters, because one generation's settings do not establish another's behavior. See the [API reference](/api/docs) for the request contract and the [migration checker](/tools/api-migration-checker) for client configuration checks.

## Know what has been tested

These checks ran through KeepRouter on September 29, 2026. They are small integration fixtures, not latency or throughput benchmarks.

| Check | Exact scope | Observed result |
|---|---|---|
| Text response | All nine IDs in the table above | Each returned a successful answer |
| Fully consumed stream | Gemini 3.8 Flash | Text, final usage and stream completion received |
| JSON-schema response | Gemini 3.5 Flash-Lite | Returned the requested `{"ok":true}` object |
| Tool conversation | Gemini 3.1 Pro Preview Custom Tools | Tool call followed by a successful continuation using its result |

For streaming, set `stream=True` and `stream_options={"include_usage": True}` and consume the stream to completion. Text chunks alone do not establish final usage. For JSON, validate the returned object in your application. For tools, retain the complete assistant message, call IDs and any non-text fields before appending the matching tool result. Validate arguments before executing a tool.

Media inputs, Gemini Live audio and Google-native tools require their own compatibility checks. The text and tool fixtures above do not establish support for those separate operations. If your application depends on them, check the selected model's documented endpoint before switching traffic.

## Calculate cost from returned usage

Use the current KeepRouter input, output and cached-input rates displayed on the exact model page. The [cost calculator](/tools/api-cost-calculator) estimates a planned workload; the console usage view records actual requests.

For a token-priced request with reported cache hits, the calculation is:

```text
cost = (input tokens - cached input tokens) / 1,000,000 × input rate
     + cached input tokens / 1,000,000 × cached-input rate
     + completion tokens / 1,000,000 × output rate
```

Use a cached-input rate only where the model publishes one and the response reports cache hits. Completion usage already includes reported reasoning tokens; do not add the reasoning detail a second time. In our Custom Tools continuation fixture, the response reported 40 input tokens, 42 completion tokens and 82 total tokens. The 35 reasoning tokens were part of those 42 completion tokens, not an extra chargeable quantity.

Compare cost per accepted result as well as cost per request. If two models differ in retry frequency or need different prompt lengths, their per-million-token prices alone do not answer which is cheaper for your task. Run the same small batch, count accepted results and divide total recorded spend by that count. Keep failed and incomplete requests visible in the comparison.

## Select the next test from the feature you need

If the application extracts fields, test schema validity and factual accuracy on several document shapes. If it uses tools, test a complete continuation. If it streams, test clean completion and cancellation. The bounded results above help choose a starting point; they do not rank the models. Google’s [Gemini 3.8 Flash model documentation](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-8-flash) describes the model, while the linked KeepRouter model page defines the customer route and price. Use the [cost calculator](/tools/api-cost-calculator) with your expected input and output before a larger paid trial.

## Move an application in a small, reversible step

Save a baseline from your current model: inputs, acceptance checks and typical usage. Run those examples against one selected Gemini ID, then compare correctness and total consumption. Include a completed stream or tool conversation if your application uses it. Keep preview models behind a configuration flag so a model change does not require rewriting the application.

After the small batch passes, increase traffic gradually and review errors, latency and spend in the console. Use [request observability](/features/api-observability) to connect a failed user action to its request record. Keep the previous model configuration available until the new workload has passed your own acceptance checks.

## Frequently asked questions

### Do Gemini requests need a Google key, and do all OpenAI endpoints work?

Use a KeepRouter key with /v1/chat/completions and the public Gemini model ID. Google authentication stays on the gateway. The current Google Cloud chat route does not expose /v1/responses, /v1/messages, generateContent or Live WebSockets.

### How do I call Gemini with the OpenAI SDK on KeepRouter?

Set base_url to https://keeprouter.com/v1, use a KeepRouter key and an exact catalog model ID. These examples use chat.completions.create.

### Should I add reasoning tokens to completion tokens?

No. Reported reasoning is already included in completion usage. Use the completion count once when applying the output rate.

### Which Gemini features were tested on KeepRouter?

On September 29, 2026: text on all nine listed IDs, streaming on 3.8 Flash, JSON on 3.5 Flash-Lite and a tool round trip on 3.1 Pro Preview Custom Tools. This is a bounded integration check.

### Where can I check the current Gemini API price?

Each KeepRouter model page shows its current customer rates. Compare the same workload and include retries when estimating cost per accepted result.

## Sources reviewed

_Article last reviewed 2026-10-03_

1. [KeepRouter API reference](https://keeprouter.com/api/docs)
2. [KeepRouter models and current prices](https://keeprouter.com/models)
3. [Gemini 3.8 Flash model card](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-8-flash)
4. [Gemini 3.1 Pro and custom tools](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-1-pro)

## Related guides

- [gemini 3.8 flash](https://keeprouter.com/models/gemini-3.8-flash.md)
- [gemini 3.5 flash lite](https://keeprouter.com/models/gemini-3.5-flash-lite.md)
- [gemini 3.1 pro preview customtools](https://keeprouter.com/models/gemini-3.1-pro-preview-customtools.md)
- [Gemini with the OpenAI SDK: compare the native, compatible and gateway paths](https://keeprouter.com/blog/gemini-openai-compatible-api-differences.md)
- [api cost calculator](https://keeprouter.com/tools/api-cost-calculator)
- [api migration checker](https://keeprouter.com/tools/api-migration-checker)

## Check this model's price and API

See the current customer rate, supported endpoint and setup example for the model discussed here.

[View model and pricing](https://keeprouter.com/models/gemini-3.8-flash)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
