What Happened
Z.ai’s official release notes date GLM-5.2 to June 16, 2026. The model focused on long-horizon tasks and extended the family to a one-million-token context window. The official repository also provides downloadable checkpoints and instructions for local serving.
Why It Matters
This release represented a material change in long-context processing, rather than only a new benchmark score. Longer agent sessions repeatedly revisit large histories of code, documents and tool outputs; reducing the cost of attention indexing helps make those sessions more practical.
Technical Details
GLM-5.2 introduced IndexShare, which reuses an attention indexer across groups of four sparse-attention layers. The developers also improved multi-token prediction for speculative decoding and exposed reasoning-effort settings. These mechanisms address serving efficiency while retaining long-context task execution. Claims about lossless context or benchmark leadership should be treated as developer evaluations, not guarantees for arbitrary workloads.