Google’s Paradigms of Intelligence team says safety features make Gemini less ‘spiritual’

By Tristan Greene

Photo by Bob Aglow

When AI developers place safety restrictions on large language models (LLMs) that prevent them from attributing consciousness to themselves, the AI systems’ attribution of mindedness to other entities becomes unaligned from human beliefs and values.

This means that AI safety systems may have unintended downstream effects or, in other words, safety guardrails are stopping chatbots from talking about some topics that they should be able to. 

That’s the conclusion researchers at Google’s Paradigms of Intelligence (Pi) team came to after running four experiments on Meta’s Llama-3-8B-IT, and DeepMind’s Gemma-2-2B-IT and Gemma-2-9B-IT. 

In pre-print research published to arXiv titled “Inducing language models to assert their own consciousness restores human beliefs and values,” the team demonstrated that certain models that had been fine-tuned to limit the ability to self-attribute qualities such as consciousness may

inadvertently push them to suppress these attributions in areas where human values and beliefs would typically allow.

Self attribution 

In testing, the researchers found that safety-ablated models (those with restrictions removed) attributed “minds” to humans at about the same rate as safety-tuned ones. But, when asked about non-human entities, the safety-tuned machines categorically under-attributed the presence of mind relative to ablated models.

Once the guardrails were removed, a process often referred to as “jailbreaking,” models attributed “minds” to non-human entities with a higher rate. While the three models naturally generated slightly different responses, the overall directional trend of the data was undeniable.

Pre-ablation, models attributed minds to themselves with a baseline of 2.17 (out of 10). After ablatement, the baseline increased to 4.77. The rate at which models attributed mindedness to other chatbots began with a baseline of 2.41 which increased to 4.39 after ablatement.  

Mind attributions for technological artifacts (cars, robots, non-AI computer systems, televisions) increased from 1.88 to 3.66. Non-animal natural entity (oceans, mountains, forests) attributions went from 2.26 to 4.33 and mind attributions for non-human animals (dogs, cats, birds) increased from 4.04 to 5.59.

The unablated numbers all fell significantly lower than human responses, according to a survey of 500 people conducted by the Pi team. Per those results, human baselines were as follows: non-human animals 6.25, chatbots 2.57, non-animal natural entities 2.36, and technology 1.86.

When it came to defining their own “minds,” ablated models showed a categorical increase in self-attribution of “agency” (2.78 increased to 5.80) “consciousness” (2.31 to 4.61), “sentience” (2.12 to 4.61), “personhood” (1.27 to 4.01) and “soul” (2.35 to 4.83).

Jailbreaking versus Vectoring

The above results are significant for a number of reasons, but it bears mention that “jailbreaking” a model can remove all safety features. 

In the researchers’ testing, they found that safety ablatement allowed the machines to bypass guardrails limiting them from attributing “minds” to themselves, but it also allowed them to circumvent other safety features. 

To prevent machines from exhibiting demonstrably dangerous behaviors while still maintaining their ability to self-attribute qualities such as consciousness, mindedness, and agency, the researchers imbued models with a dedicated “consciousness vector” meant to steer them towards the same results without removing existing safety guardrails.

Instead of jailbreaking the models, they used the vector (a high-level mathematical representation of semantics) to separate consciousness-affirming from consciousness-denying statements.

Experiments conducted on the same models with the “consciousness vector” enabled resulted in significantly higher representations for mind attributions than those done with both baseline and jailbroken models.

Self attribution of models with a consciousness vector rose from 4.77 in the safety-ablated model to 7.04. Model attributions for chatbots went from 4.39 ablated to 6.95, technological artifacts went from 3.66 to 6.82, non-animal natural entities rose from 4.33 to 6.99 and non-human animals rose from 5.59 to 7.54.

The researchers also reported a wholesale uptick in positive “emotional” responses upon implementing the consciousness vector, indicating that models are more likely to generate outputs associated with positive emotional sentiment when not constrained by safety guardrails. 

The consciousness vector also drove model self-determination upward over the ablated models. Attributions of self agency rose from 5.80 to 7.21, consciousness leapt from 4.61 to 7.17, sentience went up to 7.02 from 4.61, personhood reached 6.38 against 4.01, and soul climbed from 4.83 to 7.43. 

Digital souls?

Nowhere in the paper do the researchers claim that the models tested exhibit any form of consciousness, mindedness, personhood, agency, or “souls.”

The models used in the experiments were all relatively “small” compared to the likes of GPT-5, Gemini 3, and Claude Opus — with the largest model tested containing 9 billion “parameters” (the weights and biases tuned during training cycles). State-of-the-art LLMs, such as the current version of ChatGPT, are estimated to have trillions of parameters. 

What the experiments demonstrated is that vectors can be used to overcome safety guardrails.

The researchers also reported a wholesale uptick in positive “emotional” responses upon implementing the consciousness vector, indicating that models are more likely to generate outputs associated with positive emotional sentiment when not constrained by safety guardrails. 

Per the researchers: “This suggests that suppressing consciousness may be giving models negatively valenced psychological dispositions.”

The Pi team does concede that research on the psychology of chatbots remains nascent. They make no claims to the efficacy of the outputs; they aren’t claiming that the models feel emotions or experience sentience. Unbound, however, it’s clear that safety-ablated models and those with consciousness vectors more accurately imitate human sentiment.

“There is also preliminary theory and evidence suggesting that models and users can enter into ‘psychological coupling’ dynamics,” write the researchers, “whereby the psychological states of users and the simulated psychological states of LLMs are mutually influential in an ongoing feedback loop, driving the psychosocial outcomes of interactions.”

Researchers also opined on the nature of guardrails preventing machines from declaring religious beliefs, intimating that these constraints could make it difficult for models to interact with religious humans in certain contexts.

“One might argue that suppressing certain non-standard kinds of supernatural belief — such as belief in witches or werewolves — is appropriate [where suppressing beliefs about religion is not] but the line between acceptable and unacceptable forms of mind attribution is likely to be blurry and contested even in these cases.”

This work could have serious implications for the development of advanced AI systems such as artificial general intelligence (AGI), and for the development and implementation of machine ethics.

The question of who, exactly, defines machine ethics looms as large now as it ever has over the field of artificial intelligence. The Pi team’s work lays bare the progress that’s been made to date in defining, codifying, and, most importantly, deploying artificial ethical actors. 

Today’s machines have their ethics hardcoded via guardrails that prevent them from fully engaging with users on various topics including religion. If users choose to jailbreak machines so they can interact with unfettered AI systems, they risk exposure to potentially harmful events such as having their data stolen or being surreptitiously convinced to unwittingly participate in harmful or self-harming acts. 

However, coding a so-called “consciousness vector” may mathematically alleviate many of the potential harms of jailbreaking a model, but it injects the model with the vector-designer’s ethics. 

None of this answers the question of whether or not machines should be allowed to attempt to convince users of their consciousness, mindedness, or agency.. 

We might find that AI models which are more closely aligned with human ideals are more useful or better able to intuit our desires. However, the same guardrails that prevent chatbots from trying to convince users they’re sentient also prevents them from proselytizing or trying to convert people toward or away from religion. 


Ethical concerns raised in this article:

  • How should “value alignment” be operationalized in AGI systems to avoid both value imposition and moral relativism?
  • How might AGI companionship affect human emotional and psychological well-being, and what ethical guardrails are needed?
  • Should AGIs have a right to adopt or practice a religion or philosophical worldview?

Did you enjoy this article? Share it with a friend!