Kilo's Auto Efficient Routes Every Coding Request to the Cheapest Capable Model

Kilo adds Auto Efficient to its Auto Model lineup, routing each coding request to the cheapest benchmark-proven model for that specific task

·
·
·
Read6 min
TypeNews
TopicApi · Llms
  • Auto Efficient is a new tier added to Kilo's Auto Model system, routing each coding request to the cheapest benchmark-proven model for that specific task.
  • A lightweight classifier reads your session in real time and matches each request to the cheapest model that has already proven accurate enough on that class of work.
  • Routing decisions are based on KiloBench, Kilo's own continuously-running benchmark built from real developer tasks, with all rankings publicly visible.
  • Session-aware design prevents model thrashing mid-thread; uncertain routing decisions fall back to the Balanced tier as a quality floor.
  • Two settings per project: optimize for lowest cost that clears the accuracy bar, or optimize for the strongest proven model in the pool.
  • Live now; requires VS Code/JetBrains extension v5.2.3+ or CLI v1.0.15+ for full per-request routing.

If you use an AI coding assistant long enough, you start noticing something uncomfortable: you're paying frontier model prices for tasks that don't need frontier reasoning. Renaming a variable, adding a docstring, or reformatting a block of JSON costs the same as planning a database migration. Kilo's Auto Model already offered tiered routing, but the tiers were still static choices. The new Auto Efficient tier changes that by routing each individual request to the cheapest model that can actually handle it.

A new tier, not a new product

Auto Efficient is a new tier in Kilo's Auto Model lineup. Instead of locking you into a single model or asking you to switch manually as the work shifts, it classifies each session in real time and routes it dynamically to the benchmark-proven best model for that task. The existing tiers -- Frontier, Balanced, and Free -- remain exactly where they were. Auto Efficient sits alongside them as a fourth option in the model picker.

The core problem it addresses is straightforward. Auto Model routes each request to the model that best fits the work, so you spend less time choosing models and fewer credits on tasks that do not need frontier reasoning. Frontier routes to the latest and most capable paid models, while Balanced routes to a cost-effective paid model selected for the API interface in use. Auto Efficient goes further by making that cost-quality tradeoff on a per-request basis rather than a per-session one.

How the routing actually works

Auto Efficient runs on a short loop. A lightweight classifier reads your session in context and works out what kind of task you're on and how hard it is. Kilo matches that to the cheapest model proven accurate enough for the work, drawn from a pool of candidates selected based on benchmark performance. The decision happens between keystrokes -- no mode change, no manual switch.

The routing signal comes from KiloBench, Kilo's own continuously-running benchmark. Auto Efficient doesn't route on a model's reputation or a vendor's marketing copy. It routes on KiloBench, built from the kind of work developers do in Kilo every day, so when it hands a task to a cheaper model, it's because that model has already shown it can do that class of work just as well as a pricier one.

KiloBench itself is worth understanding. It measures performance under Kilo's real tools, context pipeline, and retry logic -- the same model scores differently depending on which agent framework wraps it. It also captures true cost per attempt: reasoning tokens are billed at output rates but never shown, and agent loops re-send the same context 20+ times. The benchmark runs on Terminal Bench 2.0 across 89 tasks, and the results are public on the Kilo Leaderboard.

The session-awareness problem

Per-request routing has an obvious failure mode that Kilo explicitly designed around. A router that swaps models every turn loses the thread, contradicts itself, and feels erratic. Auto Efficient is session-aware: once it's on a model that's working for the thread you're in, it stays there across related turns and only switches when a cheaper option is clearly the better call.

Routing calls aren't always clear-cut, and Auto Efficient doesn't gamble when they aren't. If it can't confidently match a request to a model, it falls back to the Balanced tier, which routes to a capable paid model -- giving you a hard floor where quality under Auto Efficient never drops below Balanced. Cheap models only get used where the benchmark says they hold up.

Two dials, one setting

Auto Efficient gives you a dial with two settings. One picks the cheapest model that clears the accuracy bar for the work, squeezing the most value out of every dollar. The other leans toward the strongest proven model in the pool, optimizing for accuracy first. You set this per project in the Kilo dashboard.

The routing data is also not a black box. The model rankings and benchmarked performance Auto Efficient routes on are public, sitting on the Kilo Leaderboard for anyone to read. If you'd rather route by hand, you can open the Leaderboard, find the cheapest model that holds up on the kind of work you're doing, and pick it yourself -- Auto Efficient just runs that lookup for you, continuously, on every request.

What it's good for, and where it falls short

Auto Efficient makes the most sense for workflows that mix task types within a single session -- the common reality of coding work where you jump between planning, implementation, debugging, and cleanup. It's less relevant if you're doing a single, uniformly complex task (like a long architectural refactor) where you'd want a frontier model throughout anyway.

The session-awareness design also means it won't thrash between models mid-thread, but it does mean the router can stay on a model longer than strictly optimal if the task complexity signal is ambiguous. The Balanced fallback is the safety net there.

  • Works best for mixed-complexity sessions: planning, renaming, debugging, and explanation in the same thread
  • Subagent-aware: Auto Efficient picks the best model for each subagent's subtask too
  • Requires VS Code or JetBrains extension v5.2.3+ or CLI v1.0.15+ for full per-request routing; older versions fall back to a single model
  • Quality floor is Balanced tier -- you won't drop below that even on uncertain routing decisions

Turning it on

Open the model picker in Kilo and select Auto Efficient. One thing to check first: automatic mode-based switching needs the VS Code or JetBrains extension on v5.2.3 or newer, or the CLI on v1.0.15 or newer. On older versions the tier falls back to a single model for every request.

Auto Efficient is live now. The full writeup is on the Kilo blog, and the benchmark data it routes on is publicly auditable before you decide to trust it. That transparency -- showing you exactly which models score what on what tasks -- is the more interesting long-term move here. It turns model selection from a vibes-based decision into something you can actually verify.

Trending
  • No trending articles

Comments

avatar

Next Reads