glm-5.2-fast-preview

by Z.AI

GLM-5.2-Fast-Preview is the high-speed version of Zhipu AI’s flagship model GLM-5.2, supporting a 1M ultra-long context. The model’s capabilities are aligned with the GLM-5.2 standard version, offering logical reasoning, long-text understanding, and code generation. Through inference acceleration optimizations, output TPS can reach 1.5–2× that of the GLM-5.2 standard version, significantly improving output speed. It is suitable for scenarios sensitive to output speed, such as real-time dialogue, multi-turn Agent calls, and streaming code generation.

API Pricing

Input$2.25 / 1M tokens
Output$7.89 / 1M tokens
Cache read$0.56 / 1M tokens

Specifications

Context window1,000,000 tokens
Modalitiestext
Featuresthinking, tool calling, function calling, structured outputs
Endpointschat_completions, claude_api

FAQ

What is glm-5.2-fast-preview?

GLM-5.2-Fast-Preview is the high-speed version of Zhipu AI’s flagship model GLM-5.2, supporting a 1M ultra-long context. The model’s capabilities are aligned with the GLM-5.2 standard version, offering logical reasoning, long-text understanding, and code generation. Through inference acceleration optimizations, output TPS can reach 1.5–2× that of the GLM-5.2 standard version, significantly improving output speed. It is suitable for scenarios sensitive to output speed, such as real-time dialogue, multi-turn Agent calls, and streaming code generation.

What is the context length of glm-5.2-fast-preview?

glm-5.2-fast-preview has a 1,000,000 token context window.

How much does glm-5.2-fast-preview cost?

On AIHubMix, glm-5.2-fast-preview costs $2.25 per million input tokens and $7.89 per million output tokens. Cached input reads are billed at $0.56 per million tokens.

What modalities does glm-5.2-fast-preview support?

glm-5.2-fast-preview accepts text input.

What features does glm-5.2-fast-preview support?

glm-5.2-fast-preview supports thinking, tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.

How do I call glm-5.2-fast-preview via API?

glm-5.2-fast-preview is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to glm-5.2-fast-preview — no other code changes needed.

Who develops glm-5.2-fast-preview?

glm-5.2-fast-preview is developed by Z.AI. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More from Z.AI

glm-5.2

by Z.AI

GLM-5.2 is Z.ai’s flagship model for the era of long-horizon tasks. With a truly usable…

$1.13 $0.79/1M in · $3.94 $2.76/1M out
30% off · 00:00–23:59 UTC
1,000,000 tokens context

coding-glm-5.2-free

by Z.AI

coding-glm-5.2-free is the open and free version of coding-glm-5.2. To maintain reliable…

coding-glm-5.2

by Z.AI

Currently, the special resources for this model are limited, but due to its popularity…

$0.06/1M in · $0.22/1M out

glm-5.1

by Z.AI

GLM-5.1 is Zhipu's latest flagship model, with greatly enhanced coding capabilities and…

$0.84/1M in · $3.38/1M out
200,000 tokens context

glm-image

by Z.AI

GLM-Image is Zhipu AI's new flagship image generation model. The model is trained…

$2/1M in · $2/1M out

coding-glm-5.1

by Z.AI

Only supports OpenAI-compatible formats.

$0.06/1M in · $0.22/1M out

Use glm-5.2-fast-preview via the AIHubMix unified API — one interface for every major LLM.