livenerf Catches Claude Opus 5.5 Drifting With a Daily 78-Question Trap
A new open-source benchmark called livenerf runs a frozen test panel every day for 30 days to detect whether a frontier model quietly degrades after launch.
- Developer ninjahawk released livenerf, an open-source benchmark for detecting post-launch capability drift in frontier models.
- v0 targets Claude Opus 5.5 via headless Claude Code on a Max subscription, no API key required.
- A frozen 78-question panel runs daily for 30 days with pinned CLI and exact-match graders.
- Built on Inspect with statistics from Adding Error Bars to Evals.
- Detects roughly 7.5-point accuracy swings per 10-day window using about 3.6% of a weekly plan.
- Decision rule, panel selection, and metrics are pre-registered in a public git commit.
livenerf turns model drift into a time series
Claims that a hosted frontier model has “gotten dumber” are difficult to test because users rarely have measurements from launch day. The open-source livenerf project records a fixed benchmark over time, creating evidence that can reveal post-launch changes in accuracy and output length.
Developer ninjahawk published the v0 release for Claude Opus 5.5 running through Claude Code on a Max subscription. Its design can support other models when their interfaces provide a stable, pinnable test harness.
Freeze the harness, watch the model
Hosted-model outputs remain stochastic because Claude Code does not expose sampling controls or allow reasoning to be disabled. livenerf makes each run as repeatable as the interface permits by freezing the prompts, Claude Code version, grader, execution environment, and analysis procedure.
The benchmark uses Inspect, the UK AI Security Institute’s evaluation framework. Its statistical approach follows Anthropic’s error-bar analysis, including paired comparisons and uncertainty intervals.
| Component | Implementation |
|---|---|
| Prompt panel | The same calibrated questions run on every measurement day. |
| Execution | One turn in an empty working directory, with tools, MCP servers, and memory files disabled. |
| Scoring | Exact-match grading avoids dependence on an LLM judge that could also drift. |
| Harness | The Claude Code CLI version remains pinned throughout the series. |
| Primary metric | Accuracy relative to each question’s baseline performance. |
| Secondary metric | Median output tokens per sample. |
| Storage | Raw Inspect .eval logs are retained as an append-only record. |
A 78-question panel from 2,336 candidates
The calibration pass screened 2,336 questions from GPQA Diamond, MMLU-Pro, competition mathematics, and AIME 2025-26, using four samples per question. Opus 5.5 recorded about 93% accuracy on the first sample, while 97% of questions were either correct in every run or incorrect in every run.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.