livenerf Catches Claude Opus 5.5 Drifting With a Daily 78-Question Trap

A new open-source benchmark called livenerf runs a frozen test panel every day for 30 days to detect whether a frontier model quietly degrades after launch.

·
·
·
livenerf Catches Claude Opus 5.5 Drifting With a Daily 78-Question TrapPRO
Read2 min
TypeRepo
  • Developer ninjahawk released livenerf, an open-source benchmark for detecting post-launch capability drift in frontier models.
  • v0 targets Claude Opus 5.5 via headless Claude Code on a Max subscription, no API key required.
  • A frozen 78-question panel runs daily for 30 days with pinned CLI and exact-match graders.
  • Built on Inspect with statistics from Adding Error Bars to Evals.
  • Detects roughly 7.5-point accuracy swings per 10-day window using about 3.6% of a weekly plan.
  • Decision rule, panel selection, and metrics are pre-registered in a public git commit.

livenerf turns model drift into a time series

Claims that a hosted frontier model has “gotten dumber” are difficult to test because users rarely have measurements from launch day. The open-source livenerf project records a fixed benchmark over time, creating evidence that can reveal post-launch changes in accuracy and output length.

Developer ninjahawk published the v0 release for Claude Opus 5.5 running through Claude Code on a Max subscription. Its design can support other models when their interfaces provide a stable, pinnable test harness.

Freeze the harness, watch the model

Hosted-model outputs remain stochastic because Claude Code does not expose sampling controls or allow reasoning to be disabled. livenerf makes each run as repeatable as the interface permits by freezing the prompts, Claude Code version, grader, execution environment, and analysis procedure.

The benchmark uses Inspect, the UK AI Security Institute’s evaluation framework. Its statistical approach follows Anthropic’s error-bar analysis, including paired comparisons and uncertainty intervals.

Component Implementation
Prompt panel The same calibrated questions run on every measurement day.
Execution One turn in an empty working directory, with tools, MCP servers, and memory files disabled.
Scoring Exact-match grading avoids dependence on an LLM judge that could also drift.
Harness The Claude Code CLI version remains pinned throughout the series.
Primary metric Accuracy relative to each question’s baseline performance.
Secondary metric Median output tokens per sample.
Storage Raw Inspect .eval logs are retained as an append-only record.

A 78-question panel from 2,336 candidates

The calibration pass screened 2,336 questions from GPQA Diamond, MMLU-Pro, competition mathematics, and AIME 2025-26, using four samples per question. Opus 5.5 recorded about 93% accuracy on the first sample, while 97% of questions were either correct in every run or incorrect in every run.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads