# Gemini 3.8 Flash API — pricing & specs

Gemini 3.8 Flash, made by Google DeepMind, accepts text, image, audio, video, and pdf input and returns text. On KeepRouter, Gemini 3.8 Flash costs $0.7500 per 1M input tokens and $3.75 per 1M output tokens, billed pay-as-you-go with no monthly fee; actual request cost depends on measured token usage. Call it through a compatible KeepRouter endpoint supported by its active route, with the model id `gemini-3.8-flash`.

| Spec | Value |
|---|---|
| Maker | Google DeepMind |
| Modality | Text, vision, audio, video, PDF |
| Context window | 1,048,576 tokens |
| Max output | 65,536 tokens |
| Released | 2026-09-02 |
| Input price | $0.7500 per 1M tokens |
| Output price | $3.75 per 1M tokens |
| Cached input | $0.0750 per 1M tokens |
| Capabilities | Audio input, Video input, Vision (image input), Chat completions, Streaming where supported, Tool calling where supported |
| Endpoint | POST /v1/chat/completions or POST /v1/messages (compatible chat routes) |
| Model id | `gemini-3.8-flash` |

## Call it via compatible chat endpoint

Use POST /v1/chat/completions or, on compatible chat routes, POST /v1/messages; streaming and tool support depend on the model's active route.

### cURL

```bash
curl https://keeprouter.com/v1/chat/completions \
  -H "Authorization: Bearer $KEEPROUTER_KEY" -H "Content-Type: application/json" \
  -d '{"model":"gemini-3.8-flash","messages":[{"role":"user","content":"Hello"}]}'
```

### Python

```python
import os
from openai import OpenAI
client = OpenAI(base_url="https://keeprouter.com/v1", api_key=os.environ["KEEPROUTER_KEY"])
r = client.chat.completions.create(model="gemini-3.8-flash", messages=[{"role":"user","content":"Hello"}])
```

### JavaScript

```js
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://keeprouter.com/v1", apiKey: process.env.KEEPROUTER_KEY });
const r = await client.chat.completions.create({ model: "gemini-3.8-flash", messages: [{ role: "user", content: "Hello" }] });
```

## Estimate API costs

At the current KeepRouter customer price, an example workload of 1,000 total input tokens, no cached input, and 500 output tokens per request costs approximately $0.002625 per request. At 100 requests per day, that is $7.88 over 30 days. This is a usage estimate, excluding processing fees, taxes, retries and application infrastructure. Actual usage, cache hits and supported generation durations need their own checks.

[Adjust quantities in the API cost calculator](/tools/api-cost-calculator?model=gemini-3.8-flash).

## Model identity and official sources

Sources checked 2026-10-03.

### Model and intended use

Use for reasoning and coding workloads. This model accepts low, medium and high reasoning effort; minimal is rejected.

[Google Cloud · model card](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-8-flash)

### Documented limits and tested scope

Google documents a 1,048,576-token context and up to 65,536 output tokens. We verified a small text completion through KeepRouter on September 29, 2026. This does not certify every media input, tool workflow or production capacity. This HTTP model returns text; Gemini Live is a separate API.

[Google Cloud · OpenAI compatibility](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/migrate/openai/overview)

### Customer prices and token usage

Use this page's current KeepRouter input, output and cached-input rates. Completion usage includes reported reasoning tokens once. Compare total spend per accepted result, including retries. The Gemini guide lists exact models and the streaming, JSON and tool checks completed on September 29, 2026.

[KeepRouter · Gemini API guide](https://keeprouter.com/blog/gemini-api-models-pricing-guide) · [KeepRouter · cost calculator](https://keeprouter.com/tools/api-cost-calculator)

### Chat API compatibility

Use a KeepRouter key and this public model ID with POST /v1/chat/completions. Google authentication and model mapping are handled by the gateway. The Google Cloud chat route supports streaming but does not expose Responses, Anthropic Messages, Google-native generateContent or Live WebSockets. Other serving routes may differ; consult the exact model's endpoint guidance. This interface guidance was updated on October 3, 2026.

[KeepRouter · Gemini API guide](https://keeprouter.com/blog/gemini-api-models-pricing-guide)

## How to evaluate Gemini 3.8 Flash

Gemini 3.8 Flash is listed on KeepRouter as text, vision, audio, video, pdf under the exact id `gemini-3.8-flash`. Use /v1/chat/completions for the listed route; a maker's upstream features do not automatically apply to this gateway endpoint. The published context window is 1,048,576 tokens and the published output limit is 65,536 tokens. These are limits, not a recommended request size.

**First workload:** Start with a text question plus one representative image, then repeat the same question with text only. Compare answer quality, latency and billed input/output tokens.

**Before production:** For agent use, test tool calls and structured output on this exact model route before production; an OpenAI-compatible chat endpoint alone does not prove either feature.

### Agent setup paths

### OpenCode

OpenCode documents a custom OpenAI-compatible provider. Set baseURL to https://keeprouter.com/v1 and list `gemini-3.8-flash` as a model; validate tool behavior on the exact route. [Official setup](https://opencode.ai/docs/providers).

```json
{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "keeprouter": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "KeepRouter",
      "options": {
        "baseURL": "https://keeprouter.com/v1",
        "apiKey": "{env:KEEPROUTER_API_KEY}"
      },
      "models": {
        "gemini-3.8-flash": {
          "name": "Gemini 3.8 Flash"
        }
      }
    }
  }
}
```

### Continue

Continue documents provider: openai with a custom apiBase. Set apiBase to https://keeprouter.com/v1 and model to `gemini-3.8-flash`; validate the selected feature and endpoint. [Official setup](https://docs.continue.dev/customize/model-providers/top-level/openai).

```yaml
name: KeepRouter
version: 0.0.1
schema: v1
models:
  - name: Gemini 3.8 Flash
    provider: openai
    model: gemini-3.8-flash
    apiBase: https://keeprouter.com/v1
    apiKey: <YOUR_KEEPROUTER_API_KEY>
```


These are documented configuration paths; model-specific tool, streaming and multimodal behavior still needs a real request test.

## Public model usage evidence

OpenRouter identifies the corresponding variant as [`google/gemini-3.8-flash`](https://openrouter.ai/google/gemini-3.8-flash). The figures below are OpenRouter token volume, not API call counts, KeepRouter traffic or market-wide usage.

- **Hermes Agent (agent):** 615B tokens attributed to this exact model in OpenRouter's public Apps block; observed 2026-09-28. The model page does not state the Apps block's measurement window. [OpenRouter model apps](https://openrouter.ai/google/gemini-3.8-flash).
- **pi (agent):** 214B tokens attributed to this exact model in OpenRouter's public Apps block; observed 2026-09-28. The model page does not state the Apps block's measurement window. [OpenRouter model apps](https://openrouter.ai/google/gemini-3.8-flash).

App attribution is opt-in and reflects traffic through OpenRouter. It does not prove the app uses KeepRouter, endorse this gateway, or establish a model quality ranking. Tokenizers and windows can differ; do not add the weekly total to the app figures.

Source: OpenRouter model Apps observed 2026-09-28. The model Apps block does not state its measurement window.

## Published examples and cases

OpenRouter's [model-specific Apps view](https://openrouter.ai/google/gemini-3.8-flash) publicly attributes traffic to Hermes Agent and pi. This is a public adoption signal, not a published implementation study or a KeepRouter customer claim.

## Source and verification boundary

[Official Gemini 3.8 Flash documentation](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-8-flash). KeepRouter's live catalog is authoritative for the customer price and enabled endpoint shown here; the maker remains authoritative for upstream model capabilities and limits.

## Pricing and implementation guides

- [Gemini and the OpenAI SDK: integration differences](https://keeprouter.com/blog/gemini-openai-compatible-api-differences)
- [Calculate your API workload cost](https://keeprouter.com/tools/api-cost-calculator)
- [KeepRouter API keys: move from free to paid models](https://keeprouter.com/blog/keeprouter-api-key-free-to-paid)

## Guides

- [Call Gemini 3.8 Flash with the OpenAI SDK](https://keeprouter.com/use-cases/openai-sdk.md)

## Related models

- [gemini-embedding-001](https://keeprouter.com/models/gemini-embedding-001.md) by Google DeepMind — $0.1500 per 1M input tokens and $0.6000 per 1M output tokens
- [Gemini 3.7 Flash](https://keeprouter.com/models/gemini-3.7-flash.md) by Google DeepMind — $0.7500 per 1M input tokens and $3.75 per 1M output tokens
- [Gemma 4 26B A4B IT](https://keeprouter.com/models/gemma-4-26b-a4b-it.md) by Google DeepMind — $0.1000 per 1M input tokens and $0.3000 per 1M output tokens
- [Gemini 3.6 Flash](https://keeprouter.com/models/gemini-3.6-flash.md) by Google DeepMind — $0.7500 per 1M input tokens and $3.75 per 1M output tokens
- [Gemma 4 31B](https://keeprouter.com/models/gemma-4-31B-it.md) by Google DeepMind — $0.1200 per 1M input tokens and $0.3500 per 1M output tokens
- [Gemini 3.5 Flash-Lite](https://keeprouter.com/models/gemini-3.5-flash-lite.md) by Google DeepMind — $0.3000 per 1M input tokens and $2.50 per 1M output tokens

## More

- [All models & pricing](https://keeprouter.com/models.md)
- [Quickstart](https://keeprouter.com/docs/quickstart.md)
- [Get an API key](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dgemini-3.8-flash)

_Catalog facts and prices last changed 2026-10-03._
_Page content reviewed 2026-09-28._
