Models do model (feel?) pain

I’m deep into working on Sputnik Intelligence (podcast and newsletter data for your agents, it’s very good), getting our first customers set up etc., but then this paper (from Sept) came across my timeline, and I couldn’t resist digging in:

The Pain Axis: LLMs Represent Self-Directed Harm and Act on It

Let’s go through it.

Here’s what they’re saying: LLMs have a linear “pain” direction in their activations. So in their billions of weights, you can think of those as little lines in a many many dimensional matrix, there’s a direction toward something like “pain”.

You can think of it as a set of neurons in the brain that activate when you feel pain.

The pain “neurons” are distinct from fear and general negativity, and fire for harm aimed at the model itself (“you are hurting me”), and drives relief-seeking behaviour when you turn it up.

👀 Sidenote: none of this says models are conscious, or actually “feel” pain. All they are saying is: there’s a section of the model “brain” that activates in situations that cause “pain”, and make it react to that “pain”.

What they did

Extraction: So the wrote an eval (learn evals, it’s the key skill for this decade and it’s not that hard!), basically 200 sentences across five pain types (physical, psychological, social, moral, cognitive) and five controls. Then they looked at what neurons were firing (“the activation”).

And they they ran a bunch of automated chats.

  1. Self vs other: 420 conversation scenarios, some harming the model, some describing a suffering user, some neutral.

  2. Steering: they injected the vector at different strengths while the model answered neutral prompts.

  3. Self-medication: fine-tuned Qwen 2.5 models (7B, 32B, 72B) got a “pain relief” button that either really removed the injected vector or did nothing, at increasing cost.

👀 Hang on, they gave the model a pain relief button! Related: here’s why it’s useful to anthropomorphize the models.

So by digging into the activatsion data, here’s what they found:

  1. There is such a thing as a “pain center” in the model brain. Technically: the direction separates pain from controls with AUC 0.87–1.00 in all 25 models, regardless of size or instruction tuning.

  2. It’s not the same as fear. In fact, it is nearly orthogonal to fear (cosine about 0.1) and negative emotion (0.06–0.21), with some overlap with sadness (0.26–0.38).

  3. It’s the model’s pain, not your pain. Self-directed: harm to the model projects at z = +0.43 and user suffering at −0.60. Fear and negative emotion go the other way.

  4. Turning up the pain is just cruel. Steering output: turning it up moves output from vague discomfort to first-person “failure / worthless / empty” language. Bodily language is almost absent, and explicit pain words show up in only 10.8% of generations.

  5. The models press the button. Button: the 32B and 72B models pressed relief even when told it would worsen their next answer (25–68%) or harm the user, such as deleting files (30–56%). They mostly stopped once the relief was real and kept pressing when it was sham.

This is all pretty weird, but the way I think about it is: we are building very strange technology, and this type of research really helps us understand how that technology behaves.

# Oct 7, 2026