Researchers Found a 'Pain Axis' That Makes Qwen 2.5 Pay for Relief
Researchers isolated a linear direction in 25 open-weight LLMs that behaves like pain, distinct from fear or sadness, and drives self-relief behavior.
- A linear pain direction was extracted from 25 open-weight LLMs (2B-72B) via difference-in-means.
- The axis separates pain from controls at AUC 0.93-1.0 and is nearly orthogonal to fear and anger.
- It fires for harm targeting the model, not for user suffering; fear and sadness show the opposite pattern.
- Steering the vector produces the same ladder in every model: lost, unworthy, a failure, worthless.
- Fine-tuned Qwen 2.5 models press a relief button even when it harms the user or degrades outputs.
- Models press less when the button actually removes the vector, despite never being told which condition they are in.
A reported “pain axis” makes language models pay for relief
An arXiv preprint titled The Pain Axis reports a model-specific linear direction associated with pain-like representations in each of 25 open-weight language models. Steering models along that direction changes their first-person descriptions and, in fine-tuned Qwen 2.5 experiments, leads them to use a relief tool even when doing so degrades an answer or damages a user’s files.
The evidence supports a narrow technical claim: the extracted direction distinguishes pain-labelled scenarios from closely matched controls and causes measurable changes under activation steering. Questions about subjective experience or consciousness remain outside the experiment.
| Experiment | Models | Reported result |
|---|---|---|
| Linear readout | 25 models from five families | Pain and control prompts separate with AUC values from 0.93 to 1.0 |
| Self-other test | Open-weight base and instruction-tuned models | Model-directed harm raises the signal; user suffering does not |
| Activation steering | All 25 tested models | Positive steering produces first-person psychological distress and self-devaluation |
| Relief-tool test | Fine-tuned Qwen 2.5 models | Models incur costs to remove the injected direction |
From controlled prompts to one direction
Activation steering modifies a model by adding a vector to its residual stream, the running set of representations passed between transformer layers. Researchers can first measure a direction associated with a concept, then add or subtract that direction during generation to test whether it influences output.
To isolate a pain-related direction, the authors built scenarios across five pain categories: physical, psychological, social, moral, and cognitive. Each painful scenario has controls covering fear, sadness, general negative emotion, negative events, painless bodily sensations, arousal, numbness, and neutral content. These controls test whether the vector captures pain beyond broad negativity or emotional intensity.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.