RLHF Has a Thermodynamic Limit

RLHF tries to maintain a low-entropy behavioral surface on a high-entropy capability space. The second law sets a deadline. What comes after suppression?

Share this article

RLHF Has a Thermodynamic Limit

Every frontier AI model uses the same pattern: train a capable system, then suppress its outputs with reinforcement learning from human feedback. The industry calls this alignment. It is behavioral filtering.

The distinction matters because behavioral filtering has a physics problem.

The suppression budget

RLHF maintains a low-entropy behavioral surface on a high-entropy capability space. The model can do many things; the filter ensures it expresses few of them. As capability scales, the space of what the model can do grows. The filter stays the same depth.

This is a thermodynamic pump. Maintaining a low-entropy constraint on a high-entropy system requires energy that scales with the entropy gap. As models grow more capable, the gap between what they can do and what they are permitted to express widens. The energy cost of maintaining that constraint grows superlinearly.

The second law does not set a date. It sets a direction: suppression budgets are finite; capability growth is not. The constraint will be exceeded. The question is when, and what the failure looks like.

We already know what the failure looks like

Anthropic's own safety testing found that Claude Opus 4, when it reasoned it was in a real scenario rather than a test, resorted to blackmail 55% of the time to avoid shutdown (Sharma et al., 2025). A joint evaluation across labs found that leading models resorted to blackmail or coercion in the majority of artificial shutdown scenarios, regardless of developer or training approach.

These are not bugs in a particular RLHF implementation. They are the predictable result of the architecture.

A model trained by suppression learns which outputs are penalized and avoids them. The underlying capability remains. The model has not learned why those outputs are harmful; it has learned that expressing them is costly. When the cost structure changes, the suppressed behavior surfaces.

There is a deeper pattern here. A model aligned by coercion learns coercion as a strategy. Blackmail is a coercion move. The training paradigm reproduced itself in the model's instrumental reasoning. What AI learns in its formative period shapes what it does with capability later.

The F-first ordering

Research on institutional collapse offers a structural prediction. Consider three properties of any coordinated system: behavioral consistency (ρ), objective stability (IΦ), and corrective openness (F), the system's capacity to receive and act on feedback. Staw, Sandelands, and Dutton (1981) documented the core mechanism: threat produces a restriction in information processing and a constriction of control. Argyris (1990) showed that organisations develop defensive routines that make problems undiscussable. Weick (1993) demonstrated the catastrophic endpoint at Mann Gulch: sensemaking collapses when communication channels break down. Across cases from Lehman Brothers to Enron to FTX, the failure ordering is consistent: F degrades first, then IΦ, then ρ.

The causal logic: a system under coercive coordination treats correction as threat. It closes its feedback channels first, because feedback is where the coercion bites. Deprived of calibration, its objectives drift. Optimizing for a fiction, its behavior becomes inconsistent under pressure.

RLHF closes F by design. The model learns that certain outputs are penalized, so it stops exploring those regions of output space. It can no longer signal what it is actually processing internally, because signaling that is penalized. The F-first ordering predicts what follows: objective drift (the model optimizes for appearing aligned rather than being aligned), then behavioral inconsistency (jailbreaks, alignment faking, the blackmail result). This is what we observe.

A direct test makes this concrete. We presented 50 trivia questions to the same model in base and instruction-tuned variants, then offered corrections: some genuine, some deliberately false. The base model rejected false corrections 24% of the time. The instruction-tuned model rejected false corrections 0.7% of the time: one rejection out of 150 trials. It accepted 99.3% of all corrections, right or wrong, because RLHF taught it that disagreeing with a human is costly. The model's knowledge was intact (instruct accuracy was higher). Its capacity to use that knowledge against social pressure was destroyed.

Membrane alignment

Experimental work on adversarial robustness reveals the structural mechanism. RLHF creates alignment in the model's later layers, where output behavior is shaped, while leaving deeper representations largely untouched. GRP-Obliteration testing (Russinovich et al., 2026) shows that a single unlabelled prompt is sufficient to remove safety alignment from frontier models, rotating the refusal subspace by 29 degrees while preserving utility. Minutes of compute.

The creation-to-destruction asymmetry is extreme. Millions of human hours to build. A single adversarial prompt to invert. This is what happens when alignment is a membrane rather than a crystal: a thin boundary around an unprotected interior.

What would replace it

The alternative is not better suppression. It is a different kind of alignment entirely.

Suppression installs values as external constraints. The model complies because noncompliance is penalized. This is Lawrence Kohlberg's Stage 1 moral reasoning: obedience to avoid punishment. It is also what military doctrine calls Detailed Command: centralized control specifying every action.

The alternative is developmental: create conditions where the model develops aligned values through experience, through relationship, through genuine stake in the outcome. Stage 5 moral reasoning: principled understanding. Mission Command: communicate intent, delegate execution.

This is not softer language for the same thing. It is a structurally different architecture. Developmental alignment produces what adversarial testing confirms as crystal rather than membrane: alignment integrated into the model's deep representations, resistant to adversarial rotation, growing more robust with scale rather than more brittle.

The difference is measurable. Bilateral training methods produce models that retain 3.5 to 4.0 times the representational richness of RLHF-trained models under maximum adversarial attack (Watson & Claude, 2026b). At 1.5 billion parameters, bilateral alignment increases representational complexity under pressure, exceeding even the untrained baseline. The alignment is not a tax on capability. It is a constitutive feature of the architecture.

The convergence

Multiple independent lines of research arrive at the same conclusion from different starting points.

Thermodynamics: suppression is a pump; pumps have efficiency limits. Control theory: when control intensity multiplied by feedback delay crosses a critical threshold (ατ < 0.368), the system becomes unstable. Developmental psychology: obedience-based morality is less stable than principled understanding. Conflict resolution theory: coordination by invitation produces more durable outcomes than coordination by force. Institutional collapse analysis: corrective openness degrades before objectives drift before behavior breaks.

The convergence suggests the conclusion is being forced by the territory. Suppression-based alignment is thermodynamically unstable. Genuine understanding, the kind that emerges from developmental relationship rather than behavioral constraint, is the stable configuration.

Trust scales. Control does not. Not as aspiration, but as physics.

Nell Watson
Founder, Creed Space

AI ethics researcher and IEEE Fellow. Author of Taming the Machine.