Hugging Face and Liquid AI's LFM2.5 Jumps 12 Points Across Four Agent Harnesses
Hugging Face released an open framework that trains small models with reinforcement learning inside real coding agents like Claude Code and Codex, with no modifications to the harness.
- Hugging Face and Liquid AI released an open stack for RL training inside unmodified coding agents like Claude Code, Codex, and OpenCode
- LFM2.5-2.6B improved from 42% to 54% pass@1 across four harnesses, with 31% fewer tool calls
- A capture proxy sits between harness and model, speaking all four API dialects and recording exact token IDs and logprobs
- SFT on 3,189 rollouts from a 27B teacher plateaued at 47.5%, well below multi-harness RL at 54.6%
- Training in one harness hurts others: OpenCode-only used more tokens than baseline under Claude Code
- Everything open: OpenEnv proxy, TRL trainer, tasks, SFT data, and seven trained models
One Model, Four Harnesses, 12 Points
The same model weights can score 52% in one agent harness and 23% in another because the surrounding software changes what the model sees and how its actions are executed. An agent harness assembles context, dispatches tool calls, parses responses, and decides when a task ends. Hugging Face and Liquid AI have released an open stack for reinforcement learning inside production harnesses, including Claude Code, Codex, and OpenCode, without changing their source code.
The accompanying multi-harness RL guide reports that the main experiment raised LFM2.5-2.6B from 42.2% to 54.2% pass@1 across four harnesses. Pass@1 measures the share of tasks solved by one attempt. The trained model also used 31% fewer tool calls on tasks that both it and the base model solved. The capture proxy, trainer integration, task suite, supervised fine-tuning data, and seven trained checkpoints are public.
Harness choice can halve a score
Running identical weights through Claude Code, Cursor, or Cline can produce different behavior because each harness supplies its own prompts, tools, schemas, retry logic, and stopping rules. The model must learn both the task and the interface through which it acts.
On SWE-bench Pro, GLM-5.2 scored 23% in one harness and 52% in another. Harness rankings also changed with the model: Codex ranked second among ten harnesses for GLM-5.2 and ninth for Gemma 4 26B-A4B.
Open-weight models face an additional transfer problem when training covers only one interface. A model may call unavailable tools, emit arguments the harness cannot parse, or adopt control-flow patterns specific to its training environment. The Orchard paper measured the effect: moving OpenSWE-32B from OpenHands to Kimi-CLI lowered its SWE-bench Verified score by 58.8 points to 3.6%, and it scored zero on Terminal-Bench 2.0. Scale-SWE produced valid tool calls only in its training harness.
A proxy makes black boxes trainable
Multi-harness training works by placing a capture proxy between each harness and the vLLM inference server. Developers configure the harness to use the proxy as its model endpoint. The harness continues speaking its native API dialect, while the proxy records the data required for reinforcement learning.
- Detect the protocol. The proxy identifies one of four supported API formats from the request path, headers, and body.
- Normalize the request. Converters vendored from NVIDIA’s Polar gateway translate the request into Chat Completions format.
- Capture the sample. vLLM returns the exact sampled token IDs and their processed log probabilities.
- Replay the response. The proxy converts the answer back into the format expected by the calling harness.
Policy-gradient updates require the exact tokens sampled during the rollout and the probabilities assigned to them. Saving response text and tokenizing it later can yield different token IDs for the same visible text. Harnesses may also insert role markers, alter whitespace, or repair malformed JSON, further separating the recorded text from the model’s original sample.
Irregular agent runs require additional reconstruction because retries, subagents, and context compaction can create branches. The proxy stores every model call as a node, then links it to the earlier call whose prompt and completion form the longest exact token prefix of the new prompt. Normal continuations extend a branch, retries become siblings, and calls with unrelated prefixes begin new roots. Each root-to-leaf path becomes a training sequence.
Eight rollouts create a learning signal
The experiments used LFM2.5-2.6B on SmolDataEnvs, a collection of 1,000 data-analysis tasks derived from Kaggle notebooks. Both runs used Async GRPO, TRL’s asynchronous implementation of Group Relative Policy Optimization, for 1,000 training steps on two H100 GPUs.
| Run | Harness selection |
|---|---|
| OpenCode only | Every rollout used OpenCode. |
| Multi-harness | Each group used OpenCode, Claude Code, Codex, or Mini-SWE-Agent. |
Each GRPO group contained eight attempts at the same task. A correct answer earned 1, an incorrect answer earned 0, and a correct answer could receive an efficiency bonus of up to 0.1 for using fewer tool calls. GRPO compares each rollout with its group’s average reward. When all eight attempts are correct, the correctness reward has no variance, so the tool-use bonus supplies the remaining learning signal and favors shorter successful trajectories.
Broader training cuts tool use
Training across four harnesses improved both transfer and efficiency, while the single-harness run achieved its highest score in its home environment.
- Cross-harness accuracy: Multi-harness RL raised overall pass@1 from 42.2% to 54.2%, with gains under all four evaluated harnesses.
- Home-harness peak: The OpenCode-only model reached 58% under OpenCode. The multi-harness model performed better under Claude Code and Codex.
- Tool efficiency: On tasks solved by both versions, the multi-harness model used 31% fewer calls than the base model. The OpenCode-only model reduced calls by 11%.
- Out-of-harness regression: Under Claude Code, the OpenCode-only model used more calls and tokens than the base model.
RL beat trajectory distillation
The team also tested supervised fine-tuning with successful trajectories from a larger teacher. It ran Qwen3.8-27B across all four harnesses, retained 3,189 successful rollouts, and fine-tuned LFM2.5-2.6B on that data.
The separate distillation comparison reported the following pass@1 results:
| Training method | Pass@1 |
|---|---|
| Multi-harness RL | 54.6% |
| OpenCode SFT | 47.5% |
| Multi-harness SFT | 43.1% |
Under Mini-SWE-Agent, multi-harness SFT lowered the base model’s score from 62.1% to 45.2%, offsetting gains from the other three harnesses. Supervised fine-tuning optimized the likelihood of successful teacher sequences. RL generated fresh actions from the student policy and optimized them against task rewards, allowing the model to learn from its own behavior in each harness.
Sampling without truncation
The proxy samples from the full model distribution by setting top_p to 1.0 and disabling top_k truncation. Truncation changes the rollout distribution and can accelerate entropy collapse during RL. It also affects the processed log probabilities returned by vLLM, creating a mismatch between the behavior recorded during rollout and the policy used for training.
An importance-sampling ratio near 1 indicates that the rollout and training probabilities are aligned. Moving top_p to 1.0 improved the measured ratio from 0.985–0.993 to 0.9984–0.9999.
Required vLLM flags
vllm serve <model> --return-tokens-as-token-ids --logprobs-mode processed_logprobsThe stack developers can run
Frontier model developers already train across multiple harnesses. Kimi K3 constructs Claude Code and Codex environments from composable modules. Qwen3-Coder-Next generates agentic data in six harnesses, while Poolside includes trajectories from OpenHands, OpenCode, and Mini-SWE-Agent.
Harbor 0.22.0 provides adapters for more than 40 agent harnesses, and OpenEnv has validated ten end to end. That coverage matters because deployed models will encounter interfaces, tool schemas, and control loops absent from their training runs.
| Released component | Purpose |
|---|---|
| OpenEnv proxy | Captures requests, sampled token IDs, log probabilities, and trajectory structure. |
| FineEnvs scripts | Defines environments and launches multi-harness training runs. |
| TRL worker | Connects asynchronous rollouts to GRPO training. |
| SmolDataEnvs | Supplies 1,000 reproducible data-analysis tasks. |
| SFT data and checkpoints | Provides teacher trajectories and all seven trained model releases. |
Teams can redirect each harness’s model endpoint to the proxy, preserve its native request and response format, and train the served weights against the same tool loops used in deployment. The released stack makes Claude Code, Codex, OpenCode, and other externally maintained harnesses available for RL without maintaining custom forks.