ETH Zurich's Adaptive Probes Beat DPO at Safety Without Breaking AI Transparency
Researchers train language models directly against activation probes to make them safer, more honest, and harder to jailbreak, without blinding interpretability tools.
- New paper uses activation probes as the direct training signal for alignment, replacing output-based rewards.
- Continuously updated probes cut harmfulness and boost honesty; frozen probes get trivially gamed by activation shift.
- Beats DPO and inference-time steering on safety-utility Pareto front with fewer training examples.
- Resulting models stay monitorable: fresh linear probes still hit 0.85-0.99 AUROC after fine-tuning.
- Substantially more robust against jailbreaks and abliteration attacks on open-weights models.
- Code and training scripts available at github.com/aisa-group/training_against_probes.
Adaptive probes guide safety tuning without losing decodability
Researchers at ETH Zurich and the ELLIS Institute Tübingen report that classifiers over a model’s hidden activations can provide the primary behavioral signal for alignment. Their method, called probe-guided fine-tuning, repeatedly updates those classifiers as the model changes. Across experiments involving Mistral, Llama, and Qwen models, adaptive probes reduced harmful or deceptive behavior while preserving the ability of newly trained probes to detect the same concepts.
Output-based methods such as reinforcement learning from human feedback, direct preference optimization, and constitutional AI derive supervision from generated text. A capable model may learn to satisfy an evaluator without adopting the intended behavior across contexts. Internal supervision offers another control surface, although optimization can invalidate any monitor used as a fixed target. The paper tests whether continuously refitting that monitor can prevent such evasion. The authors provide an accompanying GitHub repository.
Output scores can miss the mechanism
Most post-training pipelines reward responses that evaluators prefer. DPO learns directly from preferred and rejected response pairs, while RLHF typically uses a learned reward model. These methods can improve observable behavior, but their objectives contain limited information about the internal representations producing it.
A probe is a small classifier trained to predict a concept, such as harmfulness or dishonesty, from a model’s hidden activations. If a probe reliably separates harmful from harmless examples, its score can become a training objective: move activations associated with generated responses toward the desired side of the classifier’s boundary.
Previous work has treated this approach cautiously because a model can reduce probe loss by changing how it represents the concept. A fixed probe may then stop detecting harmful activations even when the generated behavior remains harmful. Optimization has removed the monitor’s access to the feature without resolving the underlying behavior.
A monitor that moves with the model
Probe-guided fine-tuning alternates between updating the language model and refreshing the probes that supervise it:
- The current model generates on-policy completions, meaning the samples come from the version being trained.
- The system records hidden activations from those completions.
- One or more linear or nonlinear probes score the activations for the target concept.
- A margin-based loss moves the activations toward the desired side of each probe boundary.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.