Z.ai's GLM-5.2 Beats Kimi K2.7 Code at a Tenth of Claude's Price
Kilo Code ran GLM-5.2 and Kimi K2.7 Code through the same backend task, and found the real gap is in planning, not building.

- Kilo Code benchmarked GLM-5.2 vs Kimi K2.7 Code on a two-phase plan-then-build test for a feature flag backend service.
- GLM-5.2 won the planning phase 9.0 to 8.1, making better-reasoned decisions on caching, bucketing, and API key security.
- Both models built near-identical, fully working services from GLM's plan, passing 15/15 and 14/15 checks respectively.
- A secondary comparison showed GLM-5.2's plan scored 9.0 vs Claude Fable 5's 9.1 on the same rubric, at ~1/10th the price.
- GLM-5.2 ships under an MIT license with a 1M-token context window; available now at $1.40/$4.40 per million tokens.
- The key workflow insight: a strong plan matters more than which model executes it — both models produced identical rollout results from the same spec.
Kilo Code published a head-to-head benchmark pitting two newly released open-weight coding models against each other: Z.ai's GLM-5.2 and Moonshot AI's Kimi K2.7 Code. The two are often compared as similarly priced rivals in the open-weight space, and with the latest version of each out in the same week, Kilo wanted to see how they stack up head to head. The result is a structured two-phase test that separates planning from building, and the gap between the two models showed up almost entirely in the first phase.
One task, two phases, one clear winner
Kilo ran both models through the same two-phase test. First, each model planned a backend service. They scored the plans, picked the stronger one, and then had both models build that exact plan from scratch in Kilo Code CLI. The task was a feature flag service: a backend that decides whether a feature is on for a given user and supports gradual percentage rollouts.
The task is deceptively hard. The rollout has to be deterministic: if a user is included in the first 20%, they should still be included when the rollout grows to 40%. The service also cannot solve that by storing every user assignment in a database. A weak plan waves this away. A strong plan nails the exact math.
Both nailed the hard part. Each landed on the same kind of rollout math, the kind that grows a rollout without dropping anyone already in it. But on the judgment calls the prompt left open, they diverged.
Where GLM-5.2 pulled ahead
GLM-5.2 scored 9.0 against Kimi K2.7 Code's 8.1 on the planning rubric. The difference wasn't volume of output. Kimi's plan was actually longer and included more ready-to-paste code. The gap was in which model made the hard calls and showed its reasoning.

Three specific forks separated the two plans:
- Unknown-flag caching: GLM's plan avoided unnecessary database hits by caching the "no such flag" result, but also caught the follow-up trap: if that flag gets created later, the cached negative result has to be cleared. Kimi's plan never raised that scenario.
- Rollout bucketing: GLM kept the environment out of the bucketing math and explained why: unless you deliberately change the inputs, a user should land in the same rollout slot in staging and production. Kimi included the environment in the calculation without calling out the trade-off.
- API key storage: Kimi reached for bcrypt, which is the standard answer for passwords. GLM used a single fast SHA-256 hash and explained the choice: these keys are long, random strings that cannot realistically be brute-forced, so a slow hash would add cost to every authenticated request without adding meaningful security.
In both cases, GLM made a decision and showed its reasoning. Kimi either followed the default convention or left the harder call to the builder. The plan that decides is more useful than the plan that lists, and that is why GLM's plan won.
The build phase: the plan did most of the work
For the build phase, Kilo put GLM's winning plan into an empty folder as plan.md, then gave each model a fresh Kilo Code CLI session with nothing else to work from. For Kimi, that meant building from a spec written by a rival model, without any of its own planning context carried over.

GLM passed all 15 checks. Kimi passed 14. Then came the real test: Kilo took the same 200 user IDs, evaluated them against a flag set to a 35% rollout on both finished services, and compared the answers one by one. Both services turned the flag on for the same 77 users, down to the individual IDs.
That result carries a practical implication for how you structure agentic workflows. Because the plan pinned down the exact rollout math, both builds behaved the same where it mattered. That held even in the three places where Kimi's own Round 1 plan had made a different call. When Kimi K2.7 Code built from GLM's plan, it followed GLM's decisions instead of carrying over its own earlier ones. The plan drove the implementation more than the model's habits did.
The secondary comparison: GLM-5.2 vs Claude Fable 5
With both plans scored, Kilo also pulled GLM-5.2's plan up against the one Claude Fable 5 had produced in an earlier, separate run using the same prompt and rubric. Claude Fable 5 scored 9.1, GLM-5.2 scored 9.0. Both plans made the same hard calls for the same reasons: environment kept out of the rollout hash, a fast SHA-256 for API keys, and unknown-flag lookups cached. Fable's plan was sharper in exactly one spot, since it spelled out the create-time cache trap that GLM's plan left implicit.
That near-tie matters because of the price gap behind it. Claude Fable 5 lists at $10 per million input tokens and $50 per million output. GLM-5.2 lists at $1.40 and $4.40, roughly a tenth of the price. Kilo is careful to frame this as a marker, not a verdict: one task, one run. But the direction of travel is hard to ignore.
There is also an availability angle. On June 12, 2026, a US export-control order forced Anthropic to suspend access to Claude Fable 5 and Claude Mythos 5, and because the restriction could not be enforced per user, Anthropic disabled both models for everyone. Z.ai released GLM-5.2's weights under an MIT open-source license, with documentation explicitly noting "no regional limits" and "technical access without borders." A copy of the weights you have already downloaded does not get recalled.
What this test is actually measuring
The broader finding across Kilo's last two posts is about workflow structure, not just model rankings. The sturdier finding is about the build phase: once a plan is detailed enough, the executor starts to matter less. That suggests a practical split: use the stronger model for planning, then hand the build to whichever model is cheapest or most available, and expect near-identical results.
GLM-5.2 is available now via the Z.ai API at $1.40/$4.40 per million input/output tokens. The model is available on Hugging Face, the Z.ai API, and more than 20 third-party coding environments, with a 1-million-token context window. Kilo Code also supports both models through its BYOK gateway, so you can connect your Z.ai or Kimi subscription and run the same two-phase workflow yourself. On the Artificial Analysis Intelligence Index, GLM-5.2 scores 51, leading all open-weight models including MiniMax-M3 (44) and DeepSeek V4 Pro (44).