# GLM 5.3 Flash API — pricing & specs

GLM 5.3 Flash accepts text, images, video and files and returns text with always-on reasoning. On KeepRouter, GLM 5.3 Flash costs $0.1500 per 1M input tokens and $0.5000 per 1M output tokens, billed pay-as-you-go with no monthly fee; actual request cost depends on measured token usage. Call it through a compatible KeepRouter endpoint supported by its active route, with the model id `glm-5.3-flash`.

| Spec | Value |
|---|---|
| Maker | Zhipu AI |
| Modality | Text, vision, video |
| Input price | $0.1500 per 1M tokens |
| Output price | $0.5000 per 1M tokens |
| Cached input | $0.0300 per 1M tokens |
| Capabilities | Video input, Vision (image input), Chat completions, Streaming where supported, Tool calling where supported |
| Endpoint | POST /v1/chat/completions |
| Model id | `glm-5.3-flash` |

## Call it via /v1/chat/completions

Use POST /v1/chat/completions or, on compatible chat routes, POST /v1/messages; streaming and tool support depend on the model's active route.

### cURL

```bash
curl https://keeprouter.com/v1/chat/completions \
  -H "Authorization: Bearer $KEEPROUTER_KEY" -H "Content-Type: application/json" \
  -d '{"model":"glm-5.3-flash","messages":[{"role":"user","content":"Hello"}]}'
```

### Python

```python
import os
from openai import OpenAI
client = OpenAI(base_url="https://keeprouter.com/v1", api_key=os.environ["KEEPROUTER_KEY"])
r = client.chat.completions.create(model="glm-5.3-flash", messages=[{"role":"user","content":"Hello"}])
```

### JavaScript

```js
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://keeprouter.com/v1", apiKey: process.env.KEEPROUTER_KEY });
const r = await client.chat.completions.create({ model: "glm-5.3-flash", messages: [{ role: "user", content: "Hello" }] });
```

## Estimate API costs

At the current KeepRouter customer price, an example workload of 1,000 total input tokens, no cached input, and 500 output tokens per request costs approximately $0.000400 per request. At 100 requests per day, that is $1.20 over 30 days. This is a usage estimate, excluding processing fees, taxes, retries and application infrastructure. Actual usage, cache hits and supported generation durations need their own checks.

[Adjust quantities in the API cost calculator](/tools/api-cost-calculator?model=glm-5.3-flash).

## Model identity and official sources

Sources checked 2026-10-03.

### Bounded generation checked

On October 3, 2026, this exact ID returned HTTP 200, visible text and final token usage through /v1/chat/completions. These small checks do not establish task accuracy, full-context capacity, media compatibility or sustained throughput.

[KeepRouter · API reference](https://keeprouter.com/api/docs)

### Model and evaluation task

This multimodal GLM variant accepts text, image, video and file inputs. GLM 5.3 Flash and FlashX have distinct prices; the base GLM 5.3 is text-only.

[Zhipu AI · model documentation](https://docs.z.ai/guides/vlm/glm-5.3-flash)

### API contract and migration

Reasoning remains enabled. Give the answer enough output budget. The maker states FlashX is not available on its Coding Plan; KeepRouter availability and pay-as-you-go prices do not imply entitlement on another subscription.

[Zhipu AI · model documentation](https://docs.z.ai/guides/vlm/glm-5.3-flash)

### Customer price and verification scope

Use this page's live KeepRouter USD input, output and, when shown, cached-input prices. Cached tokens are part of prompt usage and must not be counted twice. Compare spend including reasoning and retries. Reviewed documentation and catalog presence do not certify full-context capacity, media handling or throughput.

[KeepRouter · pricing and usage](https://keeprouter.com/models) · [KeepRouter · workload calculator](https://keeprouter.com/tools/api-cost-calculator)

## How to evaluate GLM 5.3 Flash

GLM 5.3 Flash is listed on KeepRouter as text, vision, video under the exact id `glm-5.3-flash`. Use /v1/chat/completions for the listed route; a maker's upstream features do not automatically apply to this gateway endpoint. Maker documentation labels the context as 1M and output as 128K. These labels are reference specifications; an exact serving token count or full-context capacity is not asserted.

**First workload:** Start with a text question plus one representative image, then repeat the same question with text only. Compare answer quality, latency and billed input/output tokens.

**Before production:** For agent use, test tool calls and structured output on this exact model route before production; an OpenAI-compatible chat endpoint alone does not prove either feature.

### Agent setup paths

### OpenCode

OpenCode documents a custom OpenAI-compatible provider. Set baseURL to https://keeprouter.com/v1 and list `glm-5.3-flash` as a model; validate tool behavior on the exact route. [Official setup](https://opencode.ai/docs/providers).

```json
{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "keeprouter": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "KeepRouter",
      "options": {
        "baseURL": "https://keeprouter.com/v1",
        "apiKey": "{env:KEEPROUTER_API_KEY}"
      },
      "models": {
        "glm-5.3-flash": {
          "name": "GLM 5.3 Flash"
        }
      }
    }
  }
}
```

### Continue

Continue documents provider: openai with a custom apiBase. Set apiBase to https://keeprouter.com/v1 and model to `glm-5.3-flash`; validate the selected feature and endpoint. [Official setup](https://docs.continue.dev/customize/model-providers/top-level/openai).

```yaml
name: KeepRouter
version: 0.0.1
schema: v1
models:
  - name: GLM 5.3 Flash
    provider: openai
    model: glm-5.3-flash
    apiBase: https://keeprouter.com/v1
    apiKey: <YOUR_KEEPROUTER_API_KEY>
```


These are documented configuration paths; model-specific tool, streaming and multimodal behavior still needs a real request test.

## Public model usage evidence

No public, exact-variant usage figure has been verified for this KeepRouter model id. Missing data is not zero usage; family-level or maker-wide traffic is not presented as this model's traffic.

## Published examples and cases

No exact-model customer case has been verified for this entry. The workload above is an evaluation recipe, not a claim of a public deployment.

## Source and verification boundary

[Official GLM 5.3 Flash documentation](https://docs.z.ai/guides/vlm/glm-5.3-flash). KeepRouter's live catalog is authoritative for the customer price and enabled endpoint shown here; the maker remains authoritative for upstream model capabilities and limits.

## Pricing and implementation guides

- [GLM 5.3, Flash and FlashX: compare inputs and costs](https://keeprouter.com/blog/glm-5-3-flash-flashx-api-guide)
- [Calculate your API workload cost](https://keeprouter.com/tools/api-cost-calculator)
- [KeepRouter API keys: move from free to paid models](https://keeprouter.com/blog/keeprouter-api-key-free-to-paid)

## Guides

- [Call GLM 5.3 Flash with the OpenAI SDK](https://keeprouter.com/use-cases/openai-sdk.md)

## Related models

- [GLM 5.3 FlashX](https://keeprouter.com/models/glm-5.3-flashx.md) by Zhipu AI — $0.3700 per 1M input tokens and $1.25 per 1M output tokens
- [GLM 5.3](https://keeprouter.com/models/glm-5.3.md) by Zhipu AI — $1.40 per 1M input tokens and $4.40 per 1M output tokens
- [GLM-5V Turbo](https://keeprouter.com/models/glm-5v-turbo.md) by Zhipu AI — $1.20 per 1M input tokens and $4.00 per 1M output tokens
- [GLM-5.2](https://keeprouter.com/models/glm-5.2.md) by Zhipu AI — $1.40 per 1M input tokens and $4.40 per 1M output tokens
- [GLM-4.5](https://keeprouter.com/models/glm-4.5.md) by Zhipu AI — $0.6000 per 1M input tokens and $2.20 per 1M output tokens
- [GLM-5.1](https://keeprouter.com/models/glm-5.1.md) by Zhipu AI — $1.40 per 1M input tokens and $4.40 per 1M output tokens

## More

- [All models & pricing](https://keeprouter.com/models.md)
- [Quickstart](https://keeprouter.com/docs/quickstart.md)
- [Get an API key](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dglm-5.3-flash)

_Catalog facts and prices last changed 2026-10-03._
_Page content reviewed 2026-10-03._
