GLM-5.3-Flash
First native multimodal model in the GLM-5 series, 320B total and 18B active parameters, low cost.
Specs
- Context
- 1.05M
- Max output
- 131.1K
- Input price
- $0.15/ 1M tokens
- Output price
- $0.5/ 1M tokens
- Released
- 2026-08-26
Context
- Context window
- 1.05M
- Max input
- —
- Max output
- 131.1K
Pricing · USD / 1M tokens
- Input
- $0.15
- Output
- $0.5
- Cache read
- $0.03
- Cache write
- —
Modalities
- Input
- TextImageVideoDocument
- Output
- Text
API
- API types
Reasoning
- Reasoning
- Yes
- Levels
- —
Info
- Status
- Active
- Released
- 2026-08-26
- Knowledge cutoff
- —
Features
- Confirmed
How to call
1 provider
zhipumodel = glm-5.3-flashreasoning = thinking.type
chat
POSThttps://api.z.ai/api/paas/v4/chat/completions
JSON
Standard format
{
"id": "glm-5.3-flash",
"object": "model",
"created": 1787702400,
"owned_by": "zhipu",
"name": "GLM-5.3-Flash",
"api": {
"types": [
"chat"
]
},
"limits": {
"context": 1048576,
"input": null,
"output": 131072
},
"modalities": {
"input": [
"text",
"image",
"video",
"document"
],
"output": [
"text"
]
},
"reasoning": {
"supported": true,
"efforts": []
},
"pricing": {
"currency": "USD",
"unit": "1M_tokens",
"input": 0.15,
"output": 0.5,
"cache_read": 0.03,
"cache_write": null
},
"features": [
"tools",
"structured_output",
"streaming",
"caching"
],
"info": {
"status": "active",
"release_date": "2026-08-26",
"knowledge_cutoff": null,
"description": "First native multimodal model in the GLM-5 series, 320B total and 18B active parameters, low cost.",
"docs": "https://docs.z.ai/guides/vlm/glm-5.3-flash",
"verified_at": "2026-10-10"
}
}Official sources
5
GLM-5.3-Flash | Z.AI Developer Documentdocs.z.ai2026-10-10Pricing | Z.AI Developer Documentdocs.z.ai2026-10-10Release Notes | Z.AI Developer Documentdocs.z.ai2026-10-10Create Chat Completion (max_tokens maximum 131072) | Z.AI Developer Documentdocs.z.ai2026-10-10GLM-4.6 config.json (max_position_embeddings 202752) | Hugging Facehuggingface.co2026-10-10