Publish time
2026/8/27Model Series
GLMInput type
Output type
Input Price
¥0.76 / 1M tokensOutput Price
¥2.66 / 1M tokensCache Write Price
¥0.95 / 1M tokensCache Read Price
¥0.2185 / 1M tokensContext Window
1,000,000Max Output Length
128,000GLM-5.3-Flash 是 GLM-5 系列首个原生多模态模型,以极致低成本架构实现超越 GLM-5.2 的更强智能: GLM-5.3-Flash 总参数量 320B、激活参数量 18B,是首个采用稀疏注意力与线性注意力混合架构的开源前沿模型。在保持精准长上下文能力的同时,显著降低计算与服务成本。相比 GLM-5.3,其注意力计算量和 KV 缓存大小分别降低 3.01 倍和 4.44 倍。 视觉能力被原生融入 Coding 循环,使模型能够主动观察界面、渲染结果与交互反馈,并据此持续测试和改进。无论是前端开发、游戏构建、Blender 3D 场景,还是 BUA、CUA 驱动的真实环境操作,模型都能在代码、浏览器和图形界面之间协同完成任务。 GLM-5.3-Flash 进一步拓展至 Office、金融研究和专业文档等工作场景。它能够自主拆解复杂目标、调用合适工具、检查并优化输出,完成从研究分析、模型构建到 PPTX、PDF、DOCX、XLSX 成品交付的完整工作流。
Zhinao API routes requests to the best-fit provider and automatically fails over to the one with highest availability.
TTFT
No data
Throughput
No data
Uptime
100.00%
Provider Model
nc/z-ai/glm-5.3-flash
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥0.56 / 1M tokens
Output Price
¥1.96 / 1M tokens
Cache Write
¥0.7 / 1M tokens
Cache Read
¥0.161 / 1M tokens
TTFT
4.86s
Throughput
5.87tps
Uptime
74.00%
Provider Model
paratera/z-ai/glm-5.3-flash
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥0.72 / 1M tokens
Output Price
¥2.52 / 1M tokens
Cache Write
¥0.9 / 1M tokens
Cache Read
¥0.207 / 1M tokens
TTFT
12.02s
Throughput
27.66tps
Uptime
100.00%
Provider Model
tokenmall-zyld/z-ai/glm-5.3-flash
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥0.48 / 1M tokens
Output Price
¥1.68 / 1M tokens
Cache Write
¥0.6 / 1M tokens
Cache Read
¥0.138 / 1M tokens
TTFT
7.64s
Throughput
29.72tps
Uptime
100.00%
Provider Model
bigmodel/z-ai/glm-5.3-flash
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥0.76 / 1M tokens
Output Price
¥2.66 / 1M tokens
Cache Write
¥0.95 / 1M tokens
Cache Read
¥0.2185 / 1M tokens
TTFT
3.82s
Throughput
50.63tps
Uptime
100.00%
Provider Model
volcengine/z-ai/glm-5.3-flash
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
128,000
Input Price
¥0.8 / 1M tokens
Output Price
¥2.8 / 1M tokens
Cache Write
¥1 / 1M tokens
Cache Read
¥0.23 / 1M tokens
Compare different providers across Zhinao API
29.77 tok/s
7.04 s
Uptime for z-ai/glm-5.3-flash across all providers
Zhinao API normalizes requests and responses across providers for you
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.360.cn/v1",
apiKey: process.env.ZHINAO_API_KEY,
});
const response = await client.chat.completions.create({
model: "z-ai/glm-5.3-flash",
messages: [
{ role: "user", content: "Hello, how are you?" }
],
temperature: 0.7,
max_tokens: 1000,
});
console.log(response.choices[0].message.content);