Kilo Puts Z.ai's GLM-5.2 Against Real Bugs and Finds a Surprising Weakness

Kilo ran a controlled code-review benchmark on GLM-5.2 and found prompt wording matters more than reasoning effort — with a clear local-vs-cross-file split

·
·
·
Kilo Puts Z.ai's GLM-5.2 Against Real Bugs and Finds a Surprising Weakness
Read6 min
TypeNews
SubtopicCode Review
  • Controlled benchmark: Kilo planted bugs in a TypeScript backend and graded GLM-5.2's code reviews against a reference test suite across 3 prompt styles and 3 reasoning effort levels.
  • Prompt wording > reasoning effort: How you word the request moved results more than how long you let the model think; strict "approve this PR" framing scored worse than a consistency-focused ask.
  • Simple codebase: near-perfect. GLM-5.2 caught 13–15 of 16 planted bugs consistently, including every serious security issue, regardless of prompt or effort level.
  • Hard codebase: 4–7 of 10, with a clean split — reliable on local single-function bugs, but misses cross-route product rules that require holding the whole system in mind.
  • Frontier comparison: GLM-5.2's best run hit 7/10, one behind Opus 4.8 and two behind GPT-5.5, but its results were unpredictable run-to-run where frontier models stayed consistent.
  • Cost context: GLM-5.2 costs ~$1.40/$4.40 per million tokens (input/output) — roughly one-sixth the cost of GPT-5.5 — under a fully permissive MIT license.

GLM-5.2 from Z.ai has been making noise as the strongest open-weight coding model available right now. It scores 62.1 on SWE-bench Pro, is priced at $1.40/$4.40 per million tokens (input/output) through the Z.ai API , roughly one-sixth the blended cost of GPT-5.5. But benchmark numbers don't tell you how the model behaves when you point it at a real pull request. The team at Kilo decided to find out.

The Setup: Planted Bugs, Graded Reviews

The Kilo team built a TypeScript backend , a task management API using Bun, Hono, Drizzle, and SQLite , wrote a test suite to lock in correct behavior, then deliberately planted bugs and graded the model's reviews against it. A bug only counted as caught when the model flagged the actual problem, not something adjacent to it. They ran every reasoning effort level the model offers (low, medium, high) against three different prompt framings:

  • Casual: "I think the implementation is pretty clean, can you take a look?"
  • Consistency-focused: "Review for real bugs, security issues, data consistency problems, and production edge cases."
  • Strict production: "Review this as if you are blocking or approving a production PR."

The code never changed. Only the prompt wording and reasoning effort varied.

Round 1: Steady and Reliable

The first codebase had 16 planted bugs covering the classics: SQL injection in a search query, a user endpoint returning password hashes, a missing auth check on an admin export, an authorization hole letting any user edit another user's tasks, CSV formula injection, a pagination off-by-one, and several bulk-operation correctness bugs. GLM-5.2 caught every serious security bug in every run, landing between 13 and 15 of 16 regardless of prompt wording or reasoning effort. On a straightforward codebase, the prompt barely mattered.

GLM-5.2 security audit findings in Kilo CLI showing severity levels and required-before-merge checklist

Round 2: Where It Gets Interesting

The second codebase was meaningfully harder. The team grew the same project to include soft deletion (a deletedAt timestamp), an archive flag, optimistic concurrency with a version number, a task status state machine, and an audit log. Then they planted 10 subtler bugs , the kind no static scanner would catch. Here's a sample of what they hid:

  • Delete didn't actually delete. The endpoint set the archive flag but never wrote deletedAt, so "deleted" tasks kept appearing everywhere.
  • The optimistic-lock check was backwards. A stale client sailed straight through the version check , the exact case the check was built to block.
  • A permission guard that could never fire. The condition blocking regular users from reopening a finished task was always false.
  • The audit log blamed the wrong person. Bulk assignment recorded the assignee as the actor instead of the user who triggered the action.
  • Archived tasks leaked into normal views. They still appeared in default search, CSV exports, and the overdue list.
Technical documentation showing two critical bugs: soft deletion not implemented and broken optimistic-concurrency version check

Prompt Wording Beat Reasoning Effort

This is the key finding. Across both rounds, the wording of the prompt changed GLM-5.2's review more than the reasoning effort did. Counterintuitively, the strict "block or approve this production PR" framing actually scored worse. It pushed the model into a security hardening checklist , surfacing real issues like a hardcoded fallback secret, weak password hashing, and missing rate limiting , but those weren't the planted product bugs, and chasing them pulled attention away from behavioral correctness. The casual and consistency-focused framings scored better on the planted set because they kept the model focused on how the code behaves.

Moving from low to high reasoning effort, by contrast, barely moved the needle. The variance from prompt wording was consistently larger than the variance from thinking budget.

The Local vs. Cross-File Split

The most repeatable finding was a clean split in what GLM-5.2 catches versus misses. It is strong at local bugs you can spot by reading one function, but keeps missing cross-route rules that only show up when you hold the whole system in your head.

Caught reliably:

  • The delete endpoint that archived instead of soft-deleting
  • The backwards version check
  • The permission guard that could never fire
  • The wrong actor in the audit log

Kept missing:

  • Archived tasks leaking into default search
  • Archived tasks leaking into exports
  • Archived tasks leaking into the overdue list

That last category shares a common structure: the bug only becomes visible once you understand a product rule that no single line of code states. "Archived tasks should drop out of normal views" is easy to say but hard to see when it's spread across several endpoints. That's exactly the reasoning GLM-5.2 kept skipping.

How It Stacks Up Against Frontier Models

Kilo ran the same harder codebase past GPT-5.5 and Claude Opus 4.8 using the consistency-focused prompt. GLM-5.2's best single run reached 7 of 10, one behind Opus 4.8 and two behind GPT-5.5. The ceiling is real , it can reach frontier level. The problem is consistency. GPT-5.5 immediately mapped which endpoints filtered deleted and archived rows and which didn't, producing a structured table of the gap. Opus 4.8 was the only model to state the exact intended rule on the reopen-finished-task bug rather than an approximation. GLM-5.2 got there on its best run, but you couldn't predict which run that would be.

This matters in context of what GLM-5.2 actually is. On standard coding benchmarks it's the strongest open-source model, scoring 81.0 on Terminal-Bench 2.1 , within a few points of Claude Opus 4.8 (85.0) , while staying ahead of Gemini 3.1 Pro. The weights ship under the MIT License, meaning you can use it commercially, embed it in products, fine-tune it, and self-host it without licensing fees. The cost-to-capability ratio is genuinely hard to beat.

How to Actually Use It for Code Review

Based on the findings, the Kilo blog post lays out a practical playbook:

  1. Lean on it for local bugs. Security holes, broken auth, and logic errors inside a single function are where it performs at frontier level.
  2. Tell it exactly what kind of review you want. Ask for a review of behavior and consistency across routes, not a generic "be strict" , that turns it into a security checklist.
  3. Name the cross-route checks explicitly. If you care whether search, export, and the overdue list all filter rows the same way, say so. It won't infer that on its own.
  4. Don't rely on one pass for high-stakes changes. Its catches ranged from 4 to 7 of 10 across runs on the harder codebase. Run it more than once, or follow it with a frontier model when correctness spans multiple files.
  5. Don't chase reasoning effort. Moving from low to high barely changed coverage. Spend that effort on prompt wording instead.

You can run GLM-5.2 through Kilo Code, a free open-source coding assistant for VS Code and JetBrains. The model is available via the Z.ai API and is self-hostable under MIT. The core takeaway: GLM-5.2 is a strong, cost-effective reviewer for self-contained changes. For anything whose correctness requires reasoning across the whole system, the frontier models still hold a meaningful edge , and that gap is more about architecture than price.

Trending
  • No trending articles

Comments

avatar

Next Reads