Back to timeline

GLM-5.2

Z.ai launches GLM-5.2 with one-million-token context and IndexShare for more efficient sparse attention.

Model Release

What Happened

Z.ai’s official release notes date GLM-5.2 to June 16, 2026. The model focused on long-horizon tasks and extended the family to a one-million-token context window. The official repository also provides downloadable checkpoints and instructions for local serving.

Why It Matters

This release represented a material change in long-context processing, rather than only a new benchmark score. Longer agent sessions repeatedly revisit large histories of code, documents and tool outputs; reducing the cost of attention indexing helps make those sessions more practical.

Technical Details

GLM-5.2 introduced IndexShare, which reuses an attention indexer across groups of four sparse-attention layers. The developers also improved multi-token prediction for speculative decoding and exposed reasoning-effort settings. These mechanisms address serving efficiency while retaining long-context task execution. Claims about lossless context or benchmark leadership should be treated as developer evaluations, not guarantees for arbitrary workloads.