Publish time
2026/8/14Model Series
GLMInput type
Output type
Input Price
¥7.6 / 1M tokensOutput Price
¥26.6 / 1M tokensCache Write Price
¥9.5 / 1M tokensCache Read Price
¥1.9 / 1M tokensContext Window
1,000,000Max Output Length
128,000GLM-5.3 是智谱最新旗舰模型,复杂软件工程与 Agent 任务能力全面进阶。它使用与 GLM-5.2 相同的基础模型——所有提升均来自后训练。与 GLM-5.2 相比,它在复杂编程和长程任务方面表现更加出色。 GLM-5.3 目前仅支持处理文本模态信息,支持 1M 上下文窗口,最大输出 Tokens 为 128K。 GLM-5.3 会始终启用思考功能,支持三个思考强度级别:low、high 和 max,并不再支持禁用思考功能。
Zhinao API routes requests to the best-fit provider and automatically fails over to the one with highest availability.
TTFT
6.95s
Throughput
31.32tps
Uptime
100.00%
Provider Model
huaweicloud/z-ai/glm-5.3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥7.2 / 1M tokens
Output Price
¥25.2 / 1M tokens
Cache Write
¥9 / 1M tokens
Cache Read
¥1.8 / 1M tokens
TTFT
4.95s
Throughput
33.91tps
Uptime
100.00%
Provider Model
nc/z-ai/glm-5.3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥5.6 / 1M tokens
Output Price
¥19.6 / 1M tokens
Cache Write
¥7 / 1M tokens
Cache Read
¥1.4 / 1M tokens
TTFT
5.21s
Throughput
2.16tps
Uptime
43.00%
Provider Model
paratera/z-ai/glm-5.3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥6.4 / 1M tokens
Output Price
¥22.4 / 1M tokens
Cache Write
¥8 / 1M tokens
Cache Read
¥1.6 / 1M tokens
TTFT
17.02s
Throughput
26.26tps
Uptime
100.00%
Provider Model
tokenmall-zyld/z-ai/glm-5.3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥4 / 1M tokens
Output Price
¥14 / 1M tokens
Cache Write
¥5 / 1M tokens
Cache Read
¥1 / 1M tokens
TTFT
3.75s
Throughput
39.82tps
Uptime
100.00%
Provider Model
bigmodel/z-ai/glm-5.3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥7.6 / 1M tokens
Output Price
¥26.6 / 1M tokens
Cache Write
¥9.5 / 1M tokens
Cache Read
¥1.9 / 1M tokens
TTFT
4.17s
Throughput
53.48tps
Uptime
100.00%
Provider Model
volcengine/z-ai/glm-5.3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥7.2 / 1M tokens
Output Price
¥25.2 / 1M tokens
Cache Write
¥9 / 1M tokens
Cache Read
¥1.8 / 1M tokens
TTFT
0.49s
Throughput
0.33tps
Uptime
100.00%
Provider Model
tencent/z-ai/glm-5.3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥7.6 / 1M tokens
Output Price
¥26.6 / 1M tokens
Cache Write
¥9.5 / 1M tokens
Cache Read
¥1.9 / 1M tokens
TTFT
4.06s
Throughput
18.91tps
Uptime
100.00%
Provider Model
alibaba/z-ai/glm-5.3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥8 / 1M tokens
Output Price
¥28 / 1M tokens
Cache Write
¥10 / 1M tokens
Cache Read
¥2 / 1M tokens
Compare different providers across Zhinao API
28.69 tok/s
8.14 s
Uptime for z-ai/glm-5.3 across all providers
Zhinao API normalizes requests and responses across providers for you
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.360.cn/v1",
apiKey: process.env.ZHINAO_API_KEY,
});
const response = await client.chat.completions.create({
model: "z-ai/glm-5.3",
messages: [
{ role: "user", content: "Hello, how are you?" }
],
temperature: 0.7,
max_tokens: 1000,
});
console.log(response.choices[0].message.content);