EN / 中文

Zhipu AI GLM-5.2: Open-Source 1M Context Model Matching Opus-Level Coding Performance

by zhidongxi·June 21, 2026

Author | Chen Junda, Editor | Mo Ying

ZhiDongXi reported on June 17 that today, Zhipu AI officially released and open-sourced its new-generation flagship model, GLM-5.2. On Code Arena, the programming evaluation system of the large model blind test platform Arena.ai, GLM-5.2 scored a high 1,595 points, ranking second on the overall leaderboard, just behind Fable 5, and ranking first among globally available models.

In the FrontierSWE benchmark, which evaluates "ultra-long-horizon, open-ended, and highly difficult software engineering tasks," GLM-5.2 currently ranks just behind Opus 4.8 and the temporarily unavailable Fable 5.

On Design Arena, which specifically evaluates model taste, GLM-5.2 achieved a global first-place performance, pushing its aesthetic capabilities to the global frontier.

On Zhihu,top contributor Toyama Nao joked that users accessing Opus through relay stations would face a new problem in the future: if Opus is actually GLM-5.2 in disguise, users might genuinely be unable to tell the difference.

Domestic and international users who experienced the actual performance of GLM-5.2 have responded enthusiastically. One developer stated bluntly, "This is the first domestic model that has reached the Opus level in my workflow."

Overseas users also reported that GLM-5.2's performance exceeded expectations, with the gap between it and Fable 5 being much smaller than anticipated. Now that Fable 5 is no longer functioning properly, overseas netizens originally thought its ban would widen the gap, but unexpectedly, GLM has almost caught up. Now it is Anthropic's turn to have a headache.

Currently, the GLM-5.2 API is live, and enterprises and users can also directly download and deploy this model on open-source platforms such as Hugging Face. Previously, ZhiDongXi had conducted in-depth reviews of Zhipu AI's models, including GLM-4.5, GLM-4.7, GLM-5, and GLM-5.1. Following the release of GLM-5.2, we immediately ran several large-scale test cases and could clearly perceive a distinct evolutionary trajectory: If GLM-4.7 achieved alignment with the top-tier programming model Sonnet 4.6 at the time, the "user experience" of GLM-5.2 is now virtually indistinguishable from Opus-level models.

In the field of AI programming models, the globally recognized top players have long been limited to Anthropic (Claude series) and OpenAI (GPT series). This time, through its first-place ranking among globally available programming models and the genuine reputation among developers as an "Opus alternative," GLM-5.2 is joining this top-tier club. It can be said that a "Big Three of Coding" landscape composed of Anthropic, OpenAI, and Zhipu AI is taking shape. At a time when closed-source giants monopolize the discourse on programming models and can revoke access permissions at any time, GLM-5.2 returns the power of choice to the vast developer community through open source.

01.4 Hours of Collaborative Programming with GLM-5.2: Nearly Exhausting the 1-Million Context Window, Fixing 16 Bugs, and Building a "Civilization" Clone from Scratch

My first hands-on task was to have GLM-5.2 develop a "Civilization"-style strategy game from scratch, iterating step-by-step from version M0 to M4. Before formal development, I first asked GLM-5.2 to write a PRD (Product Requirements Document) and discussed the specific technical implementation with it. The final technical solution was determined to use the Godot engine and GDScript to create a game with a 2.5D art style.

Version M0 served as the foundation of the entire project. In this version, GLM-5.2 consecutively created and wrote over a dozen files, generating core content such as standard map grids and basic game units. After development, GLM-5.2 quickly ran a validation check and delivered version M0.

However, this version was merely a preliminary result. The game design was still quite rough, with characters replaced only by circular icons, lacking clear game mechanics, and containing several minor interaction-level bugs.

I decided to optimize these bugs one by one at the M0 stage. Under my instructions, GLM-5.2 fixed multiple bugs, such as the information panel failing to open and initial units being unable to move. However, the fix for each bug could basically be completed within one or two rounds of dialogue, which was quite efficient. After that, I skipped version M1 and directly asked GLM-5.2 to develop version M2, which is the core of the game's depth. Without explicit requirements, GLM-5.2 autonomously judged and decided to add four major subsystems: combat system, tech tree, city economy, and resource constraints. The development workload for these new systems was substantial, and GLM-5.2 worked continuously for over 30 minutes to complete it.

During this process, GLM-5.2 strictly followed the development rules we set: complete a feature, run a test, and proceed to the next development only if there are no issues. In fact, by the later stages of this iteration, the context window had already exceeded 300,000 tokens, so it was truly commendable that GLM-5.2 could still remember the rules at this point. Version M3 transformed the game from a sandbox into a complete single match with a clear win-or-lose outcome. GLM-5.2 implemented the enemy tactical AI and expanded the map size. Although my development instructions mainly focused on iterating the game's core features, GLM-5.2 also proactively considered game optimization issues.

As the map grew larger, GLM-5.2 decided to split the terrain rendering into static and dynamic layers, and added cache optimization to the minimap, making the game run more smoothly. The work on the later M4 version mainly focused on aesthetics and playability. At this stage, GLM-5.2 demonstrated excellent aesthetics. For example, when I told it that the game's UI design "lacked a gaming feel" and was just a pile of text, it found materials on its own to update the icons and redesigned the interactive cards, elevating the visual effects of the entire game to a new level.

Finally, I encountered an unexpected bug. When the map expanded to a size of 100x100, the screen jumped violently during dragging, and I tried various methods without success. In the end, it was GLM-5.2 that successfully located the problem: it discovered that this issue had actually persisted since version M0 but only became noticeable when the map was enlarged, and it was related to UI control issues. Locating the root cause of such a problem means that GLM-5.2 can span a context length of several hundred thousand tokens to accurately pinpoint hidden bugs in the initial code. After completing all the aforementioned development tasks, we also did a simple count. In this project, GLM-5.2 used a total of 870,000 context window tokens, approaching its limit.

GLM-5.2 reviewed all the bugs it fixed during the task that approached the 1-million context length. Its statistical result was 16, which was consistent with the actual data. Meanwhile, GLM-5.2 still remembered the cause and solution of each bug, truly demonstrating reliable memory within the 1-million context scenario.

02.Reading 30 Hours of Podcast Transcripts in One Go, GLM-5.1 Falls Short

Beyond programming, the 1-million context capability of GLM-5.2 can unlock many other use cases. In daily work, I often need to process and integrate information from massive long texts, and models with larger context windows can significantly improve efficiency. In the hands-on test, I uploaded 13 AI-related podcast transcripts at once, with a total duration of over 30 hours, a text volume of about 250,000 words, which translates to at least 300,000 tokens. These podcasts are from The Lex Fridman Podcast, featuring different guests over a span of several weeks. The topics cover multiple sub-fields such as large model architecture, enterprise AI strategy, multimodality, AI safety, and the open-source ecosystem. The information is highly dispersed, with a large amount of cross-temporal echoing, supplementation, and contradiction of viewpoints. After having GLM-5.2 read all 13 transcripts at once, I assigned the following interpretation tasks:

(1) Cross-temporal Viewpoint Tracking: I asked GLM-5.2 to locate the discussion trajectory of the topic "whether the scaling law has hit a bottleneck" across all 13 transcripts. GLM-5.2 successfully identified Jensen Huang's clear denial of the "pre-training hitting a wall" theory, and also found Sam Altman's emphasis on the importance of compute during the scaling process, completely stringing together a viewpoint evolution chain spanning 30 hours of dialogue and tens of thousands of words.

GLM-5.2 ultimately provided a summary, noting that in 2023, people were still discussing single pre-training scaling, but since then, the definition of the Scaling Law has continuously expanded, evolving into four curves covering pre-training, post-training, test-time, and agents. It also judged that the current main difficulty still lies at the architectural level—whether it is truly possible to make another Transformer-level technological innovation—and found discussions on related issues by Hassabis and Terence Tao in the podcast transcripts, providing well-founded evidence.

(2) Topic Clustering: Afterward, I also asked GLM-5.2 to automatically categorize the scattered and diverse discussions into themes such as "paths to improving reasoning capabilities," "validity boundaries of synthetic data," and "mainstream choices for Agent architecture," generating consensus summaries and unresolved controversies under each theme. GLM-5.2 completed the sorting in just over 1 minute, identifying 9 major themes, each containing viewpoints from multiple different individuals, demonstrating a grasp of hundreds of thousands of tokens of context. I spot-checked several key quotes and found that GLM-5.2 basically did not hallucinate; the relevant viewpoints could all be corroborated in the podcast transcripts.

If such tasks were processed using models with conventional context windows, they could only be input in segments, summarized in batches, and then manually spliced together. The logical connections and implicit contradictions across transcripts would be lost to some extent. To verify this phenomenon, we had GLM-5.1 (with a 200,000-token context window) try the same cross-temporal viewpoint tracking task.

Ultimately, although GLM-5.1 could also read through these contents step by step, its output summary felt more like extracting and then aggregating each file after reading them one by one. GLM-5.1 failed to successfully pinpoint the details of how viewpoints changed across different periods and how they were related to each other—details that require extraction across multiple files.

However, not all tasks inevitably require the 1-million context capability of GLM-5.2. For some lightweight tasks, GLM-5.1 and GLM-5.2 do not bring about noticeable differences in user experience. For example, I had GLM-5.1 and GLM-5.2 perform the same lightweight Web UI development work, and the output speed and quality of both models were basically consistent.

For tasks such as single-file code completion, simple script writing, daily Q&A, or short document summarization, the output quality of the two models is also basically on par. The advantages of the 1-million context mainly manifest in ultra-long tasks that require cross-segment information association. For most minor tweaks in daily development, a 200K window is already sufficient; there is no need to pursue 1M just for the sake of 1M.

03.The True Challenge of 1-Million Context: Fitting It In Is Just the Beginning; Usability and Affordability Are Key

So, what technologies did Zhipu AI actually adopt on GLM-5.2 to achieve the 1-million context window and enable the model to truly utilize it effectively? In fact, Zhipu AI had already launched models with a 1-million context window during the GLM-4 era, but most of its models previously still maintained smaller context windows.
In a 1-million-level context window, merely emphasizing "length" itself has limited significance. The real challenge lies in the fact that as the context scale expands, the computational complexity of the model's attention mechanism grows quadratically. To make the 1-million-token context not just a number on a parameter sheet but truly usable, two core problems must be solved: whether the model's performance can avoid significant degradation throughout the entire process from 0 to 1 million tokens, and whether the inference cost can be controlled within a usable range. This involves a massive amount of engineering work behind the scenes.

GLM-5.2's approach to this problem is to perform collaborative optimization at both the inference infrastructure level and the model architecture level. Focusing on the efficiency bottleneck of long sequences, Zhipu AI introduced a combined solution of IndexShare, KVShare, LayerSplit, and HiSparse. At the model architecture level, Zhipu AI improved the MTP (Multi-Token Prediction) layer of GLM-5.2 to achieve better speculative decoding. They applied the combined solution of IndexShare and KVShare at the MTP layer. Previously, every prediction step in MTP required an attention calculation, whereas in multi-step MTP, GLM-5.2 only calculates the indexer at the first step. After obtaining the top-k indices, all subsequent steps directly reuse them without repeated calculations.

Among these, LayerSplit has been verified in the engineering practice of optimizing the "intelligence degradation" issue in the GLM-5 series models. The Coding Agent workloads championed by GLM are characterized by long contexts and high Prefix cache hit rates, which makes Context Parallel (CP) the main parallel strategy for Prefill nodes.

At the infrastructure level, the LayerSplit proposed by Zhipu AI has been verified in the engineering practice of optimizing the "intelligence degradation" issue in the GLM-5 series models. Targeting the characteristics of Coding Agent workloads—long contexts and high Prefix cache hit rates—this technology focuses on solving the problem of redundant KV cache storage. Its core idea is: each GPU holds only the KV Cache for a subset of layers, thereby significantly reducing the VRAM usage per card. During computation, the CP rank holding the cache for a certain layer will broadcast it to other ranks before the Attention calculation.

To further reduce overhead, Zhipu AI designed an overlapping mechanism for KV Cache broadcasting and Indexer computation, allowing the two to mask each other in time. The entire process only introduces additional Indexer Cache broadcasting equivalent to about 1/8 of the KV Cache volume, making the impact of communication costs on performance negligible. Experimental results show that within the request length range of 32k-1024k, the system throughput of GLM-5.2 achieved a 3%-192% improvement compared to GLM-5.1, with the benefits becoming more significant as the context length increases.

Meanwhile, based on the model's sparse attention characteristics, Zhipu AI also designed a hierarchical memory system named HiSparse. This system can proactively offload inactive KV cache entries to host memory, significantly alleviating GPU VRAM pressure. At the same time, it maintains a hot device cache area in the GPU HBM to store frequently accessed KV cache regions, thereby minimizing data migration overhead on the critical path. These optimizations collectively reduce the VRAM usage and latency for long-sequence inference, transforming the 1-million context from merely "runnable" to truly "affordable" and "usable." Zhipu AI stated that the online inference of GLM-5.2 relies on multiple domestic computing power platforms and has completed inference adaptation with domestic computing power platforms such as Huawei Ascend, T-Head, Moore Threads, Cambricon, Kunlunxin, Muxi, Hygon, and Biren on Day 0.

In addition, GLM-5.2 has added two new thinking effort settings: High and Max. In complex coding tasks, a higher gear can be enabled to ensure the rigor of architecture-level logic. The 1-million-level context capability of Zhipu AI's GLM-5.2 will unlock many new AI application scenarios. For example, in complex Web Search tasks, GLM-5.2 can research 12-15 mainstream K12 online programming education brands based on public materials and output a complete xlsx database, analysis reports, and charts.

Combined with Zhipu AI's Agent product AutoClaw, the 1-million context and long-horizon task capabilities of GLM-5.2 can serve white-collar scenarios such as design and legal affairs. For instance, it can write dozens of prototype pages at once, autonomously iterate and fine-tune them, and maintain brand standards and consistency in the design.

For these types of tasks, the essential difference brought by GLM-5.2 lies not in whether the result is good or bad, but in "whether it is usable or not." The scale and complexity of these tasks are unimaginable for other models that do not possess 1-million context capabilities.

04.Conclusion: Zhipu AI Completes the Technological Puzzle for Long-Horizon Tasks

Reviewing Zhipu AI's recent technological roadmap, from GLM-5.1 advancing the long-horizon task capability of open-source models to the 8-hour level, to GLM-5.2 further extending this capability with a 1M context, the trajectory of its technological puzzle is clear: first enable the model to work continuously for longer, and then equip it with a sufficiently large memory capacity. The failure of long-horizon tasks is often not because the model is not smart enough, but because it forgets the initial constraints. The 1M context solves exactly this problem. Once these capability puzzles are completed, the usability of Zhipu AI's GLM series models in real engineering tasks is expected to be further improved.

In the hands-on test, GLM-5.2 has fully run through the closed loop of understanding requirements, designing solutions, writing code, running tests, fixing bugs, and final delivery. I no longer need to break down tasks paragraph by paragraph, repeatedly feed in background information, or check whether intermediate steps deviate from the original intention. Only when a model can work for a long time and remember things can it truly possess the foundation to become a long-term collaborative partner. This is also a crucial step from "conversational AI" to "executive AI."