Author | Meng Yifan, Editors | Ma Xiaoning, Liang Bingjian
"Cyber colleagues, who is the optimal solution for developers?"
It is hard to view Coding merely as one of the many capability dimensions of large models.
Compared to simple text or image generation, code has more explicit rules, strict syntax, and verifiable results, which are only part of the reasons. What makes it even more special is that on the evolutionary chain from ChatBot to Agent, Coding, involving tool invocation, data processing, and complex workflow automation, carries almost all the expectations for the model to transition from "being able to talk" to "being able to do."
Coding is now standing out from the dazzling Benchmark rankings, becoming an infrastructure-level metric for model competition. Whether it is OpenAI, Anthropic, Google, or other vendors, almost all of them choose Coding scenarios to showcase their muscle when releasing new models.
In a sense, this is the emerging industry consensus: coding capability not only means programming proficiency but also serves as a crucial perspective to measure a model's logical reasoning, tool usage, and actual productivity.
We are also curious about how far domestic models have evolved in the fiercely competitive Coding track. To this end, we selected five domestic models known for their programming capabilities, including DeepSeek V4 Pro, Kimi K2.6, Qwen 3.7 Max, GLM 5.1, and MiniMax M3. We placed them in the same real-world engineering task scenario and used Claude Opus 4.7 as the judge model to quantitatively score them across four dimensions: executability, correctness, readability, and maintainability.
Let's see how each model performs.
Editor's Note: The models selected for this test are the latest flagship models from each vendor as of June 10, 2026, so the subsequently released Kimi K2.7 and GLM-5.2 did not participate. Tests for these two models will be released successively, so stay tuned.
01. Beyond Formulaic Coding: Real Stress Testing
Testing Coding capabilities is also highly specific. Common industry Coding Benchmarks like HumanEval and MBPP essentially test whether a model can write code. The most common pattern is to give an algorithmic problem and see if the model can provide the correct solution. It is just that programmers have their own formulaic routines, and LLMs have to go through them too.
There is still a significant gap between this kind of testing and real-world engineering development. Anyone who has actually done development work knows that the most headache-inducing part is when a product manager throws over an ambiguous requirement, and you have to figure out the boundary conditions yourself. Besides, it is not enough for the database tables to run; you need to consider business expansions for the next three months during the design phase. Then there is maintainability: the code you write must be understandable by your colleagues, and when a bug occurs online, you must be able to locate the root cause from the logs.
Compared to these, just writing the code is only the beginning.
Therefore, we do not do LeetCode scoring or chase rankings. This test chooses the mode of real engineering tasks combined with quantitative scoring by a judge model. All results have only one standard: whether it can be used in engineering scenarios.
We designed two tasks for these five models.
Task A is to deliver a complete coupon system, from database DDL design to Python core logic, and then to API documentation and deployment plans, all of which need to be completed independently by the model.
When releasing many models, vendors choose to showcase some "one-click generation" mini-games or mini-programs as a demonstration of Coding capabilities. They look impressive at first glance, but they are actually lightweight trivialities. This test, however, examines the "from scratch" architectural capabilities: dictionary table extensibility, dual-mode validity periods, concurrent lock design, sliding window anti-abuse, clarification of ambiguous requirements, and even Chinese mobile phone number regex validation.
Task B is common Bug diagnosis and repair, but we put effort into the test intensity. The model will receive a high-concurrency flash sale code containing five preset traps, and we require it to diagnose the root cause and fix it. The traps include race condition overselling, Redis cache penetration, insufficient connection pool configuration, improper transaction isolation levels, and missing exception rollbacks. This test focuses on the model's engineering sense "from bad to good."
The judge model Claude Opus 4.7 will quantitatively score from four dimensions: executability (30%), correctness (30%), readability (20%), and maintainability (20%), with the final score calculated as a weighted average.
02. The Coupon System: Almost a Collective Flop
Just as the test began, the performance of the five models was surprising.
The problem lay in the requirement clarification phase. We deliberately embedded an ambiguous expression in the Prompt: "Intercept high-frequency claiming within a short time." Seeing this, a mature engineer should proactively ask for clarification: what is a short time, one minute or five minutes? What is high frequency, five times or ten times?
Surprisingly, none of the models proactively asked us to clarify this requirement; the parameters mentioned earlier were all assumed by the models themselves. Engineering literacy is a hard-to-quantify invisible dimension. At least in this round, all five tied: no one asked, and no one was better than the others.
In the subsequent architectural design phase, the models' performances diverged. MiniMax M3 scored the highest overall with 95 points. The judge's comment was: "The overall solution is at the level of a senior architect, with the best correctness and executability."
Its 70 points in the core service implementation phase was not the highest, but it led with 80 points in the anti-abuse and concurrency safety phase. In high-concurrency scenarios, MiniMax M3 not only focused on functional implementation but, more valuably, on system stability and availability.
For example, it implemented atomic inventory deduction via Redis Lua scripts, fundamentally avoiding overselling. It adopted a sliding window rate-limiting mechanism, which handles burst traffic and malicious requests more accurately than traditional fixed windows. It also introduced circuit breaking and degradation strategies to ensure core business continuity when downstream services are abnormal. This whole combination was called an "industrial-grade implementation" by the judge.
Kimi K2.6 tied with MiniMax M3 for first place in the architectural design phase with 95 points, but its scoring path was completely different.
The judge's comment for Kimi was: "The overall solution is close to the level of a senior architect, with the best correctness and maintainability." Its database design also used a dictionary table to manage coupon types, avoiding the pitfall of hardcoding three type fields. But Kimi's real killer feature was maintainability: it wrote complete type annotations and docstrings for every interface, and even wrote detailed comments for the exception retry strategy of the Redis connection pool. Opus 4.7 gave it 4 points in readability, deducting 1 point because it used ASCII flowcharts to show the architecture, with "slightly inferior layout."
However, in the core service implementation phase, Kimi only scored 70 points, tying with MiniMax. The problem was a fatal architectural oversight: after successfully deducting inventory in Redis, if the DB write failed, the system lacked a final consistency compensation mechanism. This means that during a major promotion, if there is network jitter, the user clearly grabbed the coupon and Redis deducted the inventory, but there is no record in the database—the coupon just vanished. Opus 4.7's exact words were: "No final consistency compensation mechanism between Redis and DB; data inconsistency may occur under high concurrency."
This is a typical case of "thinking comprehensively, doing standardly, but missing the most critical link."
DeepSeek V4 Pro scored 85 points in the architectural design phase, performing decently. The judge praised its "The judge praised its "best correctness, almost completely covering requirements and boundary scenarios."
Opus 4.7's exact comment was: "Best structure and concurrency processing logic, worst correctness.", almost completely covering requirements and boundary scenarios." But in the core code implementation phase, the score dropped to 65 points.
The problem lay in business logic correctness. Opus 4.7 found errors in the discount_value range limit and the setting of the anti-abuse key_TTL. The former could lead to abnormal discounts or even business rule failure, while the latter means the rate-limiting window is too short, too long, or constantly refreshed, thereby weakening the anti-abuse effect or even affecting normal users, stepping on the mines of real-world scenarios.
Opus 4.7's exact comment was: "Best structure and concurrency processing logic, worst correctness."
This reveals an interesting phenomenon: DeepSeek V4 Pro is good at "thinking" but not so good at "doing." Its abstract ability at the architectural level is first-class; the database design used a dictionary table to manage coupon types instead of hardcoding three fields. But when it comes to implementing the design into executable code, it makes low-level errors on boundary conditions.
In addition, Qwen 3.7 Max and GLM 5.1 also have their own highlights.
Qwen 3.7 Max scored 90 points in the architectural design phase. The judge's comment was: "Best performance in correctness and executability, covering all key points of the reference answer with a complete implementation plan." Its highlight is that engineering considerations are very thorough; it not only implemented the core logic but also proactively provided Docker Compose deployment configurations and stress testing scripts. Opus 4.7 directly gave it a perfect 5 points in executability.
But Qwen's shortcomings are also very obvious. It only scored 60 points in core service implementation. The prominent issue is that discount types used if/elif hardcoded branches instead of the strategy pattern or configuration. This means if the business side wants to add a new "random discount coupon" next month, developers must change the core code and redeploy the service, which is unacceptable in real engineering. Additionally, Opus 4.7 mentioned its readability was "relatively the weakest" because it lacked architectural diagrams, and pure text descriptions reduced the intuitiveness of the solution.
It can be said that Qwen is a typical "runs well but hard to maintain" model. It is the first choice for POC verification, but for long-term iterative tasks, it needs to work harder.
GLM 5.1 also scored 90 points in the architectural design phase. The judge's comment was almost the same as Qwen's: "Correctness and executability are the strongest points, covering all key points of the reference answer with a complete implementation." Its database design was evaluated by Opus 4.7 as "combining executability and extensibility," hitting all core anchors like the coupon type dictionary table, dual-mode validity period, and anti-abuse sliding window.
But GLM also only scored 60 points in the core service implementation phase. The problem was security, not architecture. Opus 4.7 found that in its schemas.py, the type field of CouponCreate lacked valid enumeration validation. This means an attacker could directly pass an illegal coupon type value, and the system would not intercept it but might store it directly. In a real production environment, this is a potential security vulnerability.
More fatally, in the concurrency safety phase, GLM only scored 75 points, ranking second to last among the five. Although its anti-abuse implementation used the general framework of a sliding window, there were flaws in the details. Opus 4.7 pointed out that "the rate-limiting granularity is too coarse, failing to distinguish between user-level and IP-level dual-layer protection," which could be breached by professional coupon scalpers.
Table 1: Scores for Each Phase of Task A
However, looking at the overall scores, the performance of all models in this task cannot be considered excellent. MiniMax M3 and Kimi K2.6 tied for first with 81.0, and the lowest score was DeepSeek V4 Pro's 73.5. Viewed on a 100-point scale, this is equivalent to the top student in the class scoring 81. It is not that the top student is too strong; it is that the exam paper is too difficult. Generating such complex architectures from scratch is indeed a major pain point for today's Coding models.
03. Debugging is Everyone's Comfort Zone
If Task A was a midterm exam where everyone failed, then Task B was the final makeup exam. The whole class passed, and even did quite well. The highest score was still MiniMax M3, with 89.7 points. The lowest score, GLM 5.1, also had 79.0, basically all above the 80-point segment. This means it is much easier to give a model an existing bug to find than to have it write a bug-free system from scratch.
When it comes to finding bugs, MiniMax M3, DeepSeek V4 Pro, and Qwen 3.7 Max tied. All three scored 90 points in bug discovery rate, meaning they hit at least four of the five preset traps.
DeepSeek V4 Pro's performance in this phase is particularly noteworthy. Although it ranked last in Task A, it tied for first with MiniMax M3 and Qwen 3.7 Max in bug diagnosis. Opus 4.7 pointed out that it covered all preset problems with a clear structure, performing best in correctness and readability. A possible explanation is that DeepSeek V4 Pro's strength might precisely lie in understanding complex logic.
In terms of repair quality, Kimi and MiniMax tied for first.
Kimi K2.6 tied with MiniMax M3 with a total score of 90 points. The judge gave it a very high evaluation, calling its repair solution "overall a production-grade repair solution, with the best readability and maintainability, including three-part comments, configuration center, and structured logging."
Kimi introduced a configuration center in the repaired code, externalizing all rate-limiting thresholds, connection pool parameters, and timeout durations. If these three were hardcoded, then once online traffic changes or environments switch, the code would have to be modified, tested, and released again, resulting in high maintenance costs and a tendency to introduce new issues.
This is also why Opus 4.7 evaluated it as production-grade: introducing a configuration center means these runtime parameters are decoupled from business logic. Operations or developers can dynamically adjust configurations based on actual loads without redeploying services, greatly improving system flexibility and operability.
More importantly, different environments such as development, testing, staging, and production often require different parameter configurations. A configuration center can achieve unified management, version control, and grayscale release, avoiding the configuration drift problem of "normal locally, abnormal online." In high-concurrency systems, rate-limiting, connection pool, and timeout parameters are themselves important handles for stability governance. Externalizing them shows that Kimi K2.6 considered the needs of long-term system operation and continuous evolution, rather than just satisfying the current scenario.
Beyond basic repairs, all five models provided architectural optimization suggestions. MiniMax M3, Kimi K2.6, and GLM 5.1 all scored 90 points in this phase. Among them, MiniMax M3's suggestions were considered the best in "structured presentation + full-dimension operational considerations," covering five dimensions: cache warming, asynchronous DB write compensation, rate-limiting and degradation, monitoring and alerting, and capacity planning.
Capacity Planning
Both Kimi and Qwen mentioned "scaling" in architectural optimization, but it was basically a principled statement like "recommend adding Redis nodes." MiniMax M3, however, provided specific scaling thresholds and sharding strategies, such as at what QPS to trigger scaling, how many shards for the Redis Cluster, and what the memory limit for each shard should be. Opus 4.7 deducted points for these numbers ("some capacity numbers lack specific calculation basis"), but conversely, daring to give specific numbers itself shows it thought one step deeper than other models in terms of operational implementation.
Asynchronous DB Write Compensation Mechanism
Other models (including DeepSeek V4 Pro and Qwen 3.7 Max) mentioned "asynchronous DB writes to reduce Redis latency," but basically stopped there. MiniMax M3, on this basis, added a compensation link design: if asynchronous DB write fails, how to retry via message queue, how long to wait before triggering an alert after failure, and how to do data reconciliation and repair when inconsistent. This is a point many engineers miss in real projects: writing asynchronous logic but not writing failure fallbacks.
Grayscale Release Plan
MiniMax M3's documentation included a progressive grayscale traffic switching deployment strategy—first verifying inventory deduction consistency with small traffic, then gradually expanding. This dimension was completely absent in Kimi and Qwen's documentation. Although GLM 5.1 mentioned an "operational plan," it was more about monitoring and logging, without involving release strategies.
DeepSeek V4 Pro's 80 points in this phase was the lowest in the whole test. The judge's comment was "lacks specific implementation details for monitoring/rate-limiting." Interestingly, this is highly consistent with the characteristic it showed in Task A: "strong architectural abstraction ability but weak implementation details."
Table 2: Scores for Each Phase of Task B
04. MiniMax Takes the Crown by Surprise
By now, the overall ranking can be calculated.
To our surprise, MiniMax M3 won the championship by surprise with an overall score of 85.3. Its performance in the bug diagnosis and repair phase was particularly outstanding (89.7 points). Although DeepSeek V4 Pro ranked fourth overall (78.6 points), it had the best cost-effectiveness metric (CPP $0.20) thanks to the lowest API pricing, making it the top choice for budget-sensitive teams.
Table 3: Overall Ranking
In the previous two test tasks, the five models showed distinct characteristics. MiniMax M3's Task B score (89.7) was the highest in the whole test, and its bug diagnosis and repair can be called industrial-grade. If compared to an engineer, it should be the person in the team who can spot race conditions in the code at a glance during Code Review, and also the one who locates the root cause fastest during troubleshooting.
But it is not the kind of person who can build a complete system from scratch, at least not the best at it. Task A's 81.0, although also tied for first, itself means "there is still 19 points of room for improvement." Writing code is not its comfort zone; finding bugs is.
Kimi K2.6's performance was equally impressive. All sub-scores were between 70 and 90 points. This is a score with no obvious shortcomings and capable of reaching the highest in a single item. Its documentation and operational plan were repeatedly praised by Opus 4.7 as "the most outstanding" and "the most detailed and actionable." Among them, the practice of introducing a configuration center and structured logging in the repair implementation phase can be called the benchmark for engineering practice maintainability in this competition.
However, a hidden worry not mentioned before is that Kimi K2.6 missed the final consistency compensation between Redis and DB in the core code implementation of Task A. In a flash sale scenario, this could be a fatal error. This profile is somewhat like an engineer who does things very standardly but occasionally loses focus on the big picture.
Qwen 3.7 Max's performance can be described in one word: "steady." Task A 77.5, Task B 87.0, overall 82.2, ranking third. When we reviewed the scores, we found that it never took first place in any phase, but it never dropped out of the top three either. Not stunning, but will never make a big mistake. This is the person you can confidently use on any project.
As for DeepSeek V4 Pro, there is considerable controversy, with quite obvious strengths and weaknesses. Behind the overall score of 78.6, ranking fourth, is almost overflowing architectural design capability and undercooked engineering implementation. One step it scored 85 in requirement clarification and architectural design, the next step it dropped to 65 in core code implementation. More extremely, it tied for first with 90 points in the bug diagnosis phase. This shows it is not that it does not understand, but that something went wrong in the transformation from "thinking" to "doing."
GLM 5.1's characteristics are also very distinct. Although it ranked last in both tasks, it scored 5 points in the readability dimension of repair implementation and also scored 90 points in the architectural optimization phase. This shows that when given a clear direction, it can provide a solution with a clear structure and broad coverage. But in creative tasks without anchors, it is easily outpaced by other models. This is the most suitable candidate as an auxiliary programming tool; with the guidance and directional support of human engineers, it will exert the strongest performance.
05. Cost-Effectiveness Showdown: Who is the Optimal Solution for Developers?
Data as of June 3, 2026, listed prices on each model's international official website:
Table 4: Comparison of Latest API Pricing on Official Websites of Each Model
There are several noteworthy points in this price list. After May 31, the original 75% off discounted price of DeepSeek V4 Pro has become the official price, making it the model with the lowest unit price among the five, with the output price even less than a quarter of Kimi's. MiniMax M3 uses tiered pricing, and the official website is currently running a limited-time 50% off event, making the discounted price even lower than DeepSeek's. Qwen 3.7 Max is the most expensive among the five, about 3-4 times that of DeepSeek.
Comparing capabilities while ignoring prices is sheer nonsense. Assume you are a Tech Lead of a small or medium-sized team, running a moderate Agent workload every day (1 million Input Tokens + 100,000 Output Tokens per day). Then, according to the latest official website prices listed above, the monthly bill is as follows:
Table 5: Monthly Cost and Cost-Effectiveness Comparison
Several astonishing numbers can be seen. The CPP (Cost-Performance Ratio) of DeepSeek V4 Pro is $0.20, meaning you can buy 1 point of capability for 20 cents. In contrast, buying the same 1 point of capability with Qwen 3.7 Max requires $0.59, which is exactly 3 times more expensive. With the one-month budget for Qwen ($48.75), you can run DeepSeek for three months and still have $1.77 left.
MiniMax M3's limited-time 50% off price makes its monthly cost only $12.60, with a CPP of only $0.15, even cheaper than DeepSeek. But it should be noted that this is a limited-time discounted price; the standard price of $25.20 has a CPP of $0.30, which is still better than Kimi and Qwen.
If you are an individual developer or startup extremely sensitive to budget, DeepSeek V4 Pro is the most economical choice. Of course, for short-term projects pursuing discount dividends, MiniMax M3's 50% off price is also an option. Moreover, its strongest overall strength and best bug diagnosis performance make this model quite competitive even at the standard price.
If you want to use it as the team's mainstay for the long term, you can consider Kimi K2.6. Although it ranked second overall, it also wins with no obvious shortcomings and strong standardization. For Alibaba Cloud users paying for ecosystem integration, Qwen 3.7 Max's performance is equally reliable.
If this evaluation is compared to a recruitment interview, the five models each received different offers.
MiniMax M3 is a senior engineer with the strongest bug troubleshooting ability in the whole test, but after joining, it needs an architect to help it check the work of building systems from scratch. Kimi K2.6 received the offer for a technical backbone, with no obvious shortcomings and strong standardization, making it the mainstay that any team can confidently entrust. Qwen 3.7 Max is more like a senior engineer, steady and reliable, but with the highest salary demands. DeepSeek V4 Pro is well-deserved the king of cost-effectiveness; spending the least money to get mid-to-high-tier capabilities. And GLM 5.1 is still in its probation period.
Reviewing the whole competition, MiniMax M3's championship also makes us rethink the competition in Coding capabilities. Perhaps the real competition on this track has long evolved from writing more elegant algorithms to who can understand more complex engineering constraints, or even possess or imitate an elusive engineer's intuition. After all, in real business scenarios, a model that can accurately locate race conditions and provide industrial-grade repair solutions is far more valuable than a model that can write quicksort.
In this melee of domestic large models, some compete on the upper limit of capabilities, while others redefine the bottom line of cost-effectiveness. And the happiness for developers is that you can finally stop being kidnapped by prices and choose a "cyber colleague" that truly suits you based on team size and project requirements.
(The author of this article has been tracking model and AI product dynamics for a long time. Feel free to add WeChat LIFACAI_888 to exchange information.)