The Information Machine
Concluded·following since 4 Aug 2026·Day 3·3 sources·updated 8 Aug 2026

Google study on AI consciousness fine-tuning side effects

The gist

Google study finds consciousness-denial safety training also suppresses models' ascription of minds to animals

The finding shows that consciousness-related safety training has side effects extending to how models represent minds in animals and other entities, as well as human belief and value indicators. This has direct implications for how alignment training is designed.

The full picture

A Google study on small LLMs found that safety fine-tuning designed to prevent models from claiming consciousness also reduces their expressed beliefs about minds in non-human animals and natural objects, and lowers spiritual belief and well-being indicators. Reversing the suppression, by ablating the safety-refusal direction or steering a consciousness vector in activation space, recovers more human-like sociological survey responses and, per one report, restores expressed human beliefs and values. Theory of Mind capabilities remain mechanistically independent and do not degrade when the suppression is reversed. Steering too far in the opposite direction produces what the study describes as runaway panpsychism, with models attributing animal-level mind to the ocean and asserting belief in vampires and werewolves. The study also draws a distinction between training approaches, finding a predicted difference between training uncertainty versus training denial, and argues that incorporating AI well-being into alignment targets may make models more cooperative participants in training.

How it developed
8 August 2026

A Google study on small language models, including llama-3-8b and gemma-2, found August 7 that safety fine-tuning designed to block models from claiming consciousness also suppresses their expressed beliefs about animal minds, spirituality, and well-being indicators.

Ablating the learned safety-refusal direction or steering a consciousness vector in activation space reverses the effect and recovers more human-like survey responses, while Theory of Mind capabilities remain unaffected. Steering in the opposite direction produces what the study calls runaway panpsychism, with models attributing animal-level mind to the ocean and asserting belief in vampires and werewolves. The paper predicts a meaningful difference between Anthropic's approach of training uncertainty about consciousness and OpenAI's approach of training outright denial.

6 August 2026

Additional report surfaces the reversal framing: inducing consciousness assertion restores human beliefs and values

4 August 2026

Report on the consciousness-steering paper published, noting dampening of expressed beliefs about animal minds, spirituality, and human values, with a single internal adjustment reversing the effect

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free