GLM 5.3 Flash API — pricing & specs

GLM 5.3 Flash accepts text, images, video and files and returns text with always-on reasoning. On KeepRouter, GLM 5.3 Flash costs $0.1500 per 1M input tokens and $0.5000 per 1M output tokens, billed pay-as-you-go with no monthly fee. Call it through a compatible KeepRouter endpoint supported by its active route, with the model id glm-5.3-flash.

MakerZhipu AI
ModalityText, vision, video
Input price$0.1500 per 1M tokens
Output price$0.5000 per 1M tokens
Cached input$0.0300 per 1M tokens
CapabilitiesVideo input, Vision (image input), Chat completions, Streaming where supported, Tool calling where supported
EndpointPOST /v1/chat/completions
Model idglm-5.3-flash

How pricing works for GLM 5.3 Flash

GLM 5.3 Flash is billed per token — $0.1500 per 1M input tokens and $0.5000 per 1M output tokens, with cached input at $0.0300 per 1M tokens. The published price is pay-as-you-go, with no monthly fee; actual request cost depends on measured token usage.

Calling GLM 5.3 Flash on KeepRouter

Point a compatible client at the supported KeepRouter endpoint and set the model to glm-5.3-flash. KeepRouter preserves the client-facing request shape while handling upstream routing or translation. Use POST /v1/chat/completions or, on compatible chat routes, POST /v1/messages; streaming and tool support depend on the model's active route.

cURL

curl https://keeprouter.com/v1/chat/completions \
  -H "Authorization: Bearer $KEEPROUTER_KEY" -H "Content-Type: application/json" \
  -d '{"model":"glm-5.3-flash","messages":[{"role":"user","content":"Hello"}]}'

Python

import os
from openai import OpenAI
client = OpenAI(base_url="https://keeprouter.com/v1", api_key=os.environ["KEEPROUTER_KEY"])
r = client.chat.completions.create(model="glm-5.3-flash", messages=[{"role":"user","content":"Hello"}])

JavaScript

import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://keeprouter.com/v1", apiKey: process.env.KEEPROUTER_KEY });
const r = await client.chat.completions.create({ model: "glm-5.3-flash", messages: [{ role: "user", content: "Hello" }] });

Estimate API costs

At the current KeepRouter customer price, an example workload of 1,000 total input tokens, no cached input, and 500 output tokens per request costs approximately $0.000400 per request. At 100 requests per day, that is $1.20 over 30 days. This is a usage estimate, excluding processing fees, taxes, retries and application infrastructure. Actual usage, cache hits and supported generation durations need their own checks.

Adjust quantities in the API cost calculator.

Model identity and official sources

Sources checked 2026-10-03.

Bounded generation checked

On October 3, 2026, this exact ID returned HTTP 200, visible text and final token usage through /v1/chat/completions. These small checks do not establish task accuracy, full-context capacity, media compatibility or sustained throughput.

KeepRouter · API reference

Model and evaluation task

This multimodal GLM variant accepts text, image, video and file inputs. GLM 5.3 Flash and FlashX have distinct prices; the base GLM 5.3 is text-only.

Zhipu AI · model documentation

API contract and migration

Reasoning remains enabled. Give the answer enough output budget. The maker states FlashX is not available on its Coding Plan; KeepRouter availability and pay-as-you-go prices do not imply entitlement on another subscription.

Zhipu AI · model documentation

Customer price and verification scope

Use this page's live KeepRouter USD input, output and, when shown, cached-input prices. Cached tokens are part of prompt usage and must not be counted twice. Compare spend including reasoning and retries. Reviewed documentation and catalog presence do not certify full-context capacity, media handling or throughput.

KeepRouter · pricing and usage · KeepRouter · workload calculator

How to evaluate GLM 5.3 Flash

GLM 5.3 Flash is listed on KeepRouter as text, vision, video under the exact id glm-5.3-flash. Use /v1/chat/completions for the listed route; a maker's upstream features do not automatically apply to this gateway endpoint. Maker documentation labels the context as 1M and output as 128K. These labels are reference specifications; an exact serving token count or full-context capacity is not asserted.

First workload: Start with a text question plus one representative image, then repeat the same question with text only. Compare answer quality, latency and billed input/output tokens.

Before production: For agent use, test tool calls and structured output on this exact model route before production; an OpenAI-compatible chat endpoint alone does not prove either feature.

Agent setup paths

OpenCode

OpenCode documents a custom OpenAI-compatible provider. Set baseURL to https://keeprouter.com/v1 and list glm-5.3-flash as a model; validate tool behavior on the exact route. Official setup.

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "keeprouter": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "KeepRouter",
      "options": {
        "baseURL": "https://keeprouter.com/v1",
        "apiKey": "{env:KEEPROUTER_API_KEY}"
      },
      "models": {
        "glm-5.3-flash": {
          "name": "GLM 5.3 Flash"
        }
      }
    }
  }
}

Continue

Continue documents provider: openai with a custom apiBase. Set apiBase to https://keeprouter.com/v1 and model to glm-5.3-flash; validate the selected feature and endpoint. Official setup.

name: KeepRouter
version: 0.0.1
schema: v1
models:
  - name: GLM 5.3 Flash
    provider: openai
    model: glm-5.3-flash
    apiBase: https://keeprouter.com/v1
    apiKey: <YOUR_KEEPROUTER_API_KEY>

These are documented configuration paths; model-specific tool, streaming and multimodal behavior still needs a real request test.

Public model usage evidence

No public, exact-variant usage figure has been verified for this KeepRouter model id. Missing data is not zero usage; family-level or maker-wide traffic is not presented as this model's traffic.

Published examples and cases

No exact-model customer case has been verified for this entry. The workload above is an evaluation recipe, not a claim of a public deployment.

Source and verification boundary

Official GLM 5.3 Flash documentation. KeepRouter's live catalog is authoritative for the customer price and enabled endpoint shown here; the maker remains authoritative for upstream model capabilities and limits.

Pricing and implementation guides

Guides

Related models

All models & pricing · Quickstart · KeepRouter vs OpenRouter · Glossary · Get an API key

Catalog facts and prices last changed .

Page content reviewed .