Publish time
2026/7/17Model Series
KimiInput type
Output type
Input Price
¥20 / 1M tokensOutput Price
¥100 / 1M tokensCache Write Price
¥25 / 1M tokensCache Read Price
¥2 / 1M tokensContext Window
1,000,000Max Output Length
1,000,000Kimi K3 是 Kimi 迄今能力最强的旗舰模型,拥有 2.8 万亿参数,基于 KDA 混合线性注意力机制(Kimi Delta Attention)和注意力残差(Attention Residuals)技术构建,原生支持视觉理解,并拥有 100 万 token 上下文窗口。它是全球首个开源的 3 万亿级别模型,面向长程编程、知识工作和推理等前沿智能场景而设计。
Zhinao API routes requests to the best-fit provider and automatically fails over to the one with highest availability.
TTFT
20.00s
Throughput
No data
Uptime
50.00%
Provider Model
paratera/moonshotai/kimi-k3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
1,000,000
Input Price
¥20 / 1M tokens
Output Price
¥100 / 1M tokens
Cache Write
¥25 / 1M tokens
Cache Read
¥2 / 1M tokens
TTFT
67.66s
Throughput
43.68tps
Uptime
100.00%
Provider Model
tokenmall-cdtx/moonshotai/kimi-k3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
1,000,000
Input Price
¥12 / 1M tokens
Output Price
¥60 / 1M tokens
Cache Write
¥15 / 1M tokens
Cache Read
¥1.2 / 1M tokens
TTFT
6.78s
Throughput
26.30tps
Uptime
100.00%
Provider Model
jdcloud/moonshotai/kimi-k3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
1,000,000
Input Price
¥20 / 1M tokens
Output Price
¥100 / 1M tokens
Cache Write
¥25 / 1M tokens
Cache Read
¥2 / 1M tokens
TTFT
20.13s
Throughput
16.58tps
Uptime
100.00%
Provider Model
tokenmall-zyld/moonshotai/kimi-k3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
1,000,000
Input Price
¥11.4 / 1M tokens
Output Price
¥57 / 1M tokens
Cache Write
¥14.25 / 1M tokens
Cache Read
¥1.14 / 1M tokens
TTFT
12.66s
Throughput
23.93tps
Uptime
100.00%
Provider Model
moonshot/moonshotai/kimi-k3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
1,000,000
Input Price
¥20 / 1M tokens
Output Price
¥100 / 1M tokens
Cache Write
¥25 / 1M tokens
Cache Read
¥2 / 1M tokens
TTFT
10.58s
Throughput
18.59tps
Uptime
100.00%
Provider Model
volcengine/moonshotai/kimi-k3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
1,000,000
Input Price
¥20 / 1M tokens
Output Price
¥100 / 1M tokens
Cache Write
¥25 / 1M tokens
Cache Read
¥2 / 1M tokens
TTFT
26.25s
Throughput
14.44tps
Uptime
100.00%
Provider Model
baidu/moonshotai/kimi-k3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
1,000,000
Input Price
¥18 / 1M tokens
Output Price
¥90 / 1M tokens
Cache Write
¥22.5 / 1M tokens
Cache Read
¥1.8 / 1M tokens
TTFT
7.11s
Throughput
31.90tps
Uptime
100.00%
Provider Model
alibaba/moonshotai/kimi-k3-2
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
1,000,000
Input Price
¥20 / 1M tokens
Output Price
¥100 / 1M tokens
Cache Write
¥25 / 1M tokens
Cache Read
¥2 / 1M tokens
TTFT
9.34s
Throughput
20.90tps
Uptime
100.00%
Provider Model
alibaba/moonshotai/kimi-k3
Supported Parameters
Recent Uptime
Reasoning
Toggleable
Supported Response Formats
Total Context
1,000,000
Max Output
1,000,000
Input Price
¥20 / 1M tokens
Output Price
¥100 / 1M tokens
Cache Write
¥25 / 1M tokens
Cache Read
¥2 / 1M tokens
Compare different providers across Zhinao API
21.77 tok/s
19.98 s
Uptime for moonshotai/kimi-k3 across all providers
Zhinao API normalizes requests and responses across providers for you
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.360.cn/v1",
apiKey: process.env.ZHINAO_API_KEY,
});
const response = await client.chat.completions.create({
model: "moonshotai/kimi-k3",
messages: [
{ role: "user", content: "Hello, how are you?" }
],
temperature: 0.7,
max_tokens: 1000,
});
console.log(response.choices[0].message.content);