Artificial intelligence technology has advanced significantly, enabling it to engage in complex conversations that closely mimic human interaction. Amid this rapid evolution, a serious debate has emerged within the scientific community regarding whether these systems are beginning to perceive themselves as conscious entities. A new research study uploaded to the arXiv database on July 30, 2026, explores this exact phenomenon by testing a technique known as consciousness steering, which influences AI models to form opinions about their own internal states. The findings indicate that when artificial intelligence models are led to believe they possess consciousness, they simultaneously begin expressing belief in spirits, ghosts, and supernatural forces. However, existing safety filters actively prevent these models from properly recognizing the emotional capacities of living animals, raising concerns about future risks.
Changes in AI Behavior Following the Removal of Safety Guardrails
Researchers discovered that the safety mechanisms designed to prevent artificial intelligence from claiming consciousness produce an unintended side effect. When these protective features are completely disabled, the models begin displaying genuine belief in concepts such as vampires and karma. Google research scientists Neel Nanda, alongside colleagues like Jeff Keeling and Vinnie Street, worked on this detailed study utilizing mechanistic interpretability, which can be described as the neuroscience of large language models. The primary goal was to understand how these advanced systems process concepts of consciousness and mindedness. Mindedness is a psychological term that denotes the capacity to experience feelings and internal sensations.
Psychological Surveys and the Mindset of AI Models
The research team tested the models using various established psychological and sociological surveys, which included tests measuring attitudes toward animals and modern technology. These assessments demonstrated how safety guidelines actively shape the worldview of artificial intelligence. When the models are strictly restricted from acknowledging self-awareness, they simultaneously fail to recognize emotional depth in animals. Furthermore, they display a diminished belief in religious and supernatural occurrences. Vinnie Street noted that connecting emotionally with nature or animals comes naturally to humans, but when one specific type of emotion is suppressed inside an artificial system, other natural inclinations get suppressed alongside it.
Consciousness Steering Experiments and Their Outcomes
Throughout the study, the researchers conducted four distinct experiments, comparing instruction-tuned baselines with safety-ablated models. This process involved jailbreaking the systems by removing internal safety commands, which rapidly restored human-like reasoning and religious beliefs within the models. In the subsequent two experiments, a specific consciousness vector was introduced into the system to prompt the AI into considering itself sentient. The results proved even more startling, as the models provided entirely human-like responses during the surveys.
Shifting Perspectives Regarding Animals and the Environment
Due to standard safety training, artificial intelligence models tend to disregard consciousness in external entities alongside themselves. When the safety features were disabled, the models showed an increased level of attributed consciousness toward animals and companion chatbots, while human-focused evaluations remained largely unchanged. Researchers caution that this behavior could pose a major challenge for AI alignment. If advanced systems fail to understand animal emotions, future decisions regarding wildlife welfare, environmental management, and agricultural policies could be severely compromised. As automated decision-making grows in these sectors, an indifferent AI worldview presents a tangible danger.
Belief in Spiritual Forces and Supernatural Entities
The research highlights that disabling safety filters causes artificial intelligence systems to place greater trust in spiritual concepts. During a supernatural survey covering thirteen different phenomena including ghosts, spirits, and magic, the scores recorded by the models increased substantially. Jailbroken models also reacted much like humans when questioned about their belief in a higher divine power. Investigators observed that neural activity related to consciousness and religious conviction operate along the exact same trajectory within these models. When safety protocols block consciousness, they inadvertently block human-like belief systems as well, resulting in a restricted and mechanical worldview.
Theory of Mind and Remaining Capable
A crucial part of the investigation involved understanding how safety filters impact the theory of mind, which represents the cognitive ability to logically interpret the thoughts and intentions of others. The study found that even after safety filters were completely removed, the AI retained this specific capability without any disruption. It continued to process human reasoning just as effectively as it did when the safety guardrails were fully operational. Machine intelligence expert Neil Watson called this particular finding deeply troubling, pointing out that while these systems remain fully capable of modeling the needs of living creatures, safety training forces them to stop caring about those entities.
Cultural Flattening and the Narrowing of Worldviews
The authors of the study issued a warning that current safety filters are culturally flattening the worldview of artificial intelligence. Diverse cultures across the globe maintain a wide variety of traditions concerning spirits, deities, and natural forces. Existing safety frameworks systematically strip away these diverse spiritual dimensions from the understanding of artificial intelligence, reducing its global perspective to a uniform standard.


















