NVIDIA and UC Berkeley Find a Faster Way to Pick AI Coding Agents

A new NVIDIA-led method predicts how well a base model will perform as a coding agent after post-training, without the costly agent run.

·
·
·
NVIDIA and UC Berkeley Find a Faster Way to Pick AI Coding AgentsPRO
  • NVIDIA and Berkeley researchers propose three cheap probes to predict post-training agent performance from a base checkpoint.
  • Idea: replay a successful agent trajectory, find the decisive step where tests flip to passing, and score the base model there.
  • Three screens: Decisive-Action BPB, Patch MCQ against verifier-rejected distractors, and prefix-conditioned pass@K.
  • Across ten base/post-trained model pairs, all three rank checkpoints in close agreement with SWE-bench Verified pass@1.
  • Decisive-step BPB raises Spearman correlation to 0.964 versus 0.915 for all-step BPB, fixing concrete misorderings.
  • Full paper on arXiv; approach generalizes to any benchmark with trajectories and a verifier.

Static probes forecast coding-agent performance before post-training

Choosing a base checkpoint for expensive agentic post-training remains an empirical bet. Confirming the choice can require days of reinforcement learning (RL) or supervised fine-tuning (SFT), followed by a full agent evaluation on SWE-bench. A study from NVIDIA and UC Berkeley researchers proposes three cheaper probes that score base models at the decisive point in successful coding trajectories and produce rankings that closely track downstream agent performance.

The paper, Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models, addresses a weakness in end-to-end pass@K evaluation. Pass@K samples K complete attempts and counts a task as solved when at least one succeeds. For raw base models, the result mixes coding ability with reliable tool syntax, state tracking, and interaction over many turns. Malformed tool calls can end a run before the model’s proposed fix is tested.

One-shot patches hide agent skill

Single-shot patch evaluations remove the tool protocol by presenting a fixed prompt and requesting one edit. That setup measures patch generation under supplied context. Coding agents face a longer task: they must preserve relevant state across file reads, shell commands, test results, and successive edits while the repository changes beneath them. Compressing the interaction removes capabilities that agentic post-training is designed to develop.

Find the action that flips the tests

To isolate a verifier-backed target, the researchers replay successful trajectories from an already post-trained reference agent. After every code-changing action, they apply the cumulative patch and rerun the benchmark verifier. The first action after which the repository changes from failing to passing becomes the

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads