Intelligent Internet's Zenith Jumps GPT-5.5 From 5th to 1st on Hardest Coding Benchmark
Intelligent Internet's open-source Zenith harness lifts GPT-5.5 from 5th to 1st on FrontierSWE, beating Claude Fable without needing a bigger model

- Zenith open-sourced: Intelligent Internet releases Zenith, an agent harness that lifts GPT-5.5 from 5th to 1st on FrontierSWE, beating Claude Fable.
- Same model, better harness: GPT-5.5 scores 2.06 avg rank with Zenith vs. 5.53 on its default Codex harness — no model swap needed.
- Four subsystems: Deep Planning, Adaptive Workers, Deep Testing (evidence-traced), and Adaptive Skills form the core control loop.
- Meta-Zenith automates harness construction: A meta-level system generates task-specific Zenith harnesses from a sketch and feedback loop.
- Context: Fable 5 suspended: Claude Fable 5 was pulled by US export controls, making open harnesses around accessible models newly critical.
- Cost-efficient: Zenith achieves best mean rank at $175/task vs. $407/task for the brute-force RALPH baseline.
The conventional wisdom in AI agent work is simple: when your agent gets stuck, upgrade the model. Intelligent Internet's Zenith just ran a controlled experiment that challenges that assumption directly. By wrapping GPT-5.5 in a purpose-built agent harness, they moved the same model from 5th place to 1st on FrontierSWE , the hardest public long-horizon software engineering benchmark , beating Claude Fable in the process. The harness is now open source.
The benchmark that breaks every agent
FrontierSWE, built by Proximal AI, is not a typical coding benchmark. It contains ultra-long-horizon, open-ended technical challenges such as optimizing compilers or training state-of-the-art models for protein prediction. Agents are given 20 hours per task; despite this, most models barely make progress on any task, making FrontierSWE one of the few unsaturated public benchmarks. On average, agents run for 11 hours per task and fail to solve almost all of them.
The benchmark exposes a specific failure mode that Intelligent Internet calls premature completion: agents don't give up, they declare victory too early. The tests they write for themselves are superficial enough to make a wrong solution look right. The agent submits, the independent test suite fails, and nobody catches it.
Fifth to first, same model
Intelligent Internet ran the full 17-task FrontierSWE suite with Zenith wrapping GPT-5.5. The results are stark: by mean@5 (average rank across five runs), Zenith scored 2.06 average rank with 92% dominance, landing first. The same GPT-5.5 model on its default Codex harness scored 5.53, sitting fifth. The benchmark, the trial budget, and the base model were identical. Only the control loop changed.
The gap is widest on implementation tasks , the longest-horizon work on the benchmark. GPT-5.5 under Codex ranks 7.40 there. With Zenith, it ranks 1.60, ahead of every other entry including Fable. That's the result Intelligent Internet cares most about, because Anthropic's own launch notes for Fable said the longer and more complex the task, the larger Fable's lead grows.

How the harness actually works
Zenith is built around four subsystems controlled by a single LLM orchestrator that decides what to do at every milestone boundary:
- Deep Planning , reads the task spec and produces a milestone-level strategy that can be revised mid-run. If the executor discovers that a language's allocator model breaks an assumption the plan made, the orchestrator replans rather than continuing on a broken blueprint.
- Adaptive Workers , execute the plan feature-by-feature. The orchestrator can spawn a new worker when a milestone needs a capability the existing workers lack, or retire one that has finished its scope.
- Deep Testing , validates at every milestone boundary through two layers: a build/test/lint gate, and an evidence-traced review against the task's acceptance criteria. A milestone seals only when both pass.
- Adaptive Skills , capture techniques that worked and register them for reuse. When a worker discovers a pattern for debugging a tricky compile-time failure, the orchestrator lifts it into a reusable skill. Later milestones that hit the same error get the fix automatically.
The key word is adaptive. Zenith is not recursive self-improvement , the model weights are frozen and the agent doesn't rewrite its own code. What changes during a run is the shape of work: which workers exist, what testing layers are active, what skills have been learned, and how the plan is structured. The orchestrator's decision trace is fully readable and replayable.
Meta-Zenith: automating the harness itself
The first Zenith harness was built by hand , manually written system prompts, worker definitions, validation rules, and orchestration logic. That doesn't scale. Each new task family needs different worker roles, validation strategies, and stopping policies. A harness that works for one class of long-running tasks may not transfer cleanly to another.
Meta-Zenith automates this construction process. Given a simplified orchestrator-worker sketch, a task specification, and feedback from a training loop, Meta-Zenith produces a fully configured Zenith harness. The output includes:
- A Contract: the task's goals, acceptance criteria, and constraints
- A Milestone graph: ordered by dependency and test coverage
- Worker specs: task-specific roles, prompts, tools, and skills
- Validator specs: build gates and fidelity checks
- Skill seeds: reusable patterns likely to recur
- A Stop/replan policy: when to continue, revise, add workers, or stop
To train Meta-Zenith, Intelligent Internet continuously collects their own internal engineering and research tasks and runs the harness-construction loop over that stream. This keeps the optimization process separate from the benchmark itself , the final FrontierSWE evaluation used finalized harnesses, not per-task manual tuning.
Why this matters right now
The timing of this release is not accidental. Anthropic was forced to disable all access to its newest AI models, Fable 5 and Mythos 5, after the U.S. Commerce Department used national security export controls to bar the company from distributing the models to any foreign national. Given the scope of the directive, Anthropic argued it had no choice but to disable the models for all users, though access to its less powerful Claude models including Opus 4.8 was not affected. Meanwhile, OpenAI's GPT-5.6 launched as a limited preview for a small group of trusted partners.
Developers who had built on Fable 5 focused on reliability, and many treated the episode as an argument for open-weight or self-hosted models that cannot be cut off from the outside. Zenith is a direct answer to that concern: an open-source harness that squeezes frontier performance out of models you can actually run, rather than depending on gated access to the most capable ones.
The internal benchmark results from the technical report are worth noting too. Zenith achieved the best mean rank across 8 long-horizon tasks while costing less than half of the brute-force RALPH baseline:
| Method | Mean Rank (lower is better) | Mean Cost per Task | Task Wins |
|---|---|---|---|
| One-session | 5.00 | $22.21 | 0 |
| Plan-RALPH | 4.00 | $161.53 | 0 |
| Milestone-RALPH | 2.88 | $209.47 | 0 |
| RALPH | 1.75 | $407.58 | 3 |
| Zenith | 1.38 | $175.68 | 5 |
What it's useful for
Zenith is built for tasks that take hours, not seconds. The practical use-cases it targets:
- Optimizing compilers or rendering pipelines end-to-end
- Training models against hidden benchmarks
- Building production-grade services (e.g., a PostgreSQL-compatible server backed by SQLite)
- Performance engineering across large codebases
- Applied ML research tasks that require iterative experimentation
It's less suited to short-horizon tasks where a single-session agent already works well. The overhead of milestone planning, adaptive workers, and evidence-traced validation pays off only when the task is genuinely long and the failure mode is premature completion rather than inability to start.
Getting started
The Zenith repository is public now, with task configs, model and version settings, per-run budgets, run logs, evaluator outputs, and full results including cost, token, and runtime summaries. The repo includes a working example , an Angry Birds-style physics puzzle game built end-to-end by Zenith as one of the eight benchmarked tasks. More details on Meta-Zenith and additional open-source releases are coming soon, with GLM-5.2 integration teased as the next update.
The deeper implication here is a field-level assumption that needs updating: the model is not the only lever. When the strongest models are gated by export controls or limited previews, the harness around the model you can run is where your real leverage is. Zenith makes that lever open source.