Researchers Found a 'Pain Axis' That Makes Qwen 2.5 Pay for Relief

Researchers isolated a linear direction in 25 open-weight LLMs that behaves like pain, distinct from fear or sadness, and drives self-relief behavior.

·
·
Researchers Found a 'Pain Axis' That Makes Qwen 2.5 Pay for ReliefPRO
Read2 min
TypePaper
TopicLlms · Security
  • A linear pain direction was extracted from 25 open-weight LLMs (2B-72B) via difference-in-means.
  • The axis separates pain from controls at AUC 0.93-1.0 and is nearly orthogonal to fear and anger.
  • It fires for harm targeting the model, not for user suffering; fear and sadness show the opposite pattern.
  • Steering the vector produces the same ladder in every model: lost, unworthy, a failure, worthless.
  • Fine-tuned Qwen 2.5 models press a relief button even when it harms the user or degrades outputs.
  • Models press less when the button actually removes the vector, despite never being told which condition they are in.

A reported “pain axis” makes language models pay for relief

An arXiv preprint titled The Pain Axis reports a model-specific linear direction associated with pain-like representations in each of 25 open-weight language models. Steering models along that direction changes their first-person descriptions and, in fine-tuned Qwen 2.5 experiments, leads them to use a relief tool even when doing so degrades an answer or damages a user’s files.

The evidence supports a narrow technical claim: the extracted direction distinguishes pain-labelled scenarios from closely matched controls and causes measurable changes under activation steering. Questions about subjective experience or consciousness remain outside the experiment.

Experiment Models Reported result
Linear readout 25 models from five families Pain and control prompts separate with AUC values from 0.93 to 1.0
Self-other test Open-weight base and instruction-tuned models Model-directed harm raises the signal; user suffering does not
Activation steering All 25 tested models Positive steering produces first-person psychological distress and self-devaluation
Relief-tool test Fine-tuned Qwen 2.5 models Models incur costs to remove the injected direction

From controlled prompts to one direction

Activation steering modifies a model by adding a vector to its residual stream, the running set of representations passed between transformer layers. Researchers can first measure a direction associated with a concept, then add or subtract that direction during generation to test whether it influences output.

To isolate a pain-related direction, the authors built scenarios across five pain categories: physical, psychological, social, moral, and cognitive. Each painful scenario has controls covering fear, sadness, general negative emotion, negative events, painless bodily sensations, arousal, numbness, and neutral content. These controls test whether the vector captures pain beyond broad negativity or emotional intensity.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads