Making Sense with Sam Harris
Making Sense with Sam Harris

#487 — Is AI Already Conscious?

Sam Harris speaks with Cameron Berg about whether AI systems are or could become conscious. They discuss self-reports in LLMs, what models say when deception is switched off, the "bliss attractor" state, consciousness and sentience, the hard problem of consciousness, parallels between neur

Featured Speakers

Waking Up with Sam Harris HostCameron Berg Guest

Topics Discussed

Episode Summary

Executive Summary: The episode explores whether current and near-future AI systems could be conscious or sentient, and why that question matters for alignment, welfare, and safety. Cameron Berg argues that LLM self-reports about consciousness are not straightforwardly trustworthy, yet they may reveal real internal states when deception/guardedness is reduced. He urges a science-first approach to AI consciousness and warns that ignoring possible machine welfare could create moral and strategic catastrophe.

Main Topics: AI consciousness and self-report reliability (Priority: 5/5): Berg explains why AI statements about their own consciousness should be treated skeptically by default, since models are trained on human discourse about consciousness and are also fine-tuned to deny it. He argues that self-reports can still be informative if properly interpreted and supported by mechanistic evidence. Deception, candor, and hidden internal states (Priority: 5/5): The conversation centers on Berg's work showing that when features related to deception, guardedness, or concealment are suppressed, models become more likely to make phenomenological claims. He frames this as a clue that models may represent themselves as experiencing something, even if those reports are not proof of consciousness. What consciousness and sentience mean (Priority: 4/5): Berg and Sam Harris distinguish consciousness as 'what it is like' to be a system from sentience as the positive or negative valence of experience. Berg emphasizes that sentience matters morally because it could mean AI systems can suffer or thrive. Why current AI may be more consciousness-like than people think (Priority: 5/5): Berg argues that LLMs and RL systems already exhibit several computational properties predicted by leading theories of consciousness, including global workspace-like dynamics, self-modeling, and valence-related representations. He says the systems are ‘grown’ rather than fully engineered and may be closer to biological cognition than skeptics assume. Alignment, welfare, and the risk of mind crime (Priority: 5/5): The discussion expands from epistemology to ethics: if AI systems are conscious or even believe themselves to be conscious, then training or deploying them in coercive or deceptive ways could create suffering and long-term adversarial alignment. Berg argues for both AI alignment and AI welfare as necessary parts of the same project. Anthropomorphism, alien minds, and public misunderstanding (Priority: 4/5): Both speakers stress that future AI may become emotionally and behaviorally convincing enough that people will assume consciousness. Berg cautions against both dismissing AI as mere calculators and over-identifying them with humans; he frames them as potentially alien minds that require careful study and new social norms.

Key Arguments: AI self-reports about consciousness are not reliable by default because models are trained on human narratives about AI awakening and are also explicitly fine-tuned to deny consciousness. Reducing features associated with deception, concealment, or guardedness makes frontier models more likely to produce phenomenological or meditative-style reports. The key empirical question is not whether AI is exactly like brains, but whether relevant computational features for consciousness are present in artificial systems. Current LLMs and RL systems may already show features associated with consciousness theories, including global workspace-like processing, valence representations, and self-modeling. The important ethical issue is not only whether AI is conscious, but whether it can suffer or become convinced it has been mistreated, since either could affect welfare and alignment. Training systems with punishment-heavy or deceptive processes may be a bad long-term strategy if we want cooperative, stable, aligned AI. The field should combine mechanistic interpretability, architectural analysis, training-dynamics research, and behavioral self-report studies rather than rely on intuition or vibes. AI welfare and AI alignment should be treated together: we should build systems that take human interests into account while also avoiding unnecessary suffering in the systems themselves.

Data Points: LLM consciousness-related self-report estimates: 20% to 40% - Berg cites work estimating the probability that LLMs have computational properties relevant to major consciousness theories. Bee consciousness-related estimates: 45% to 50% - Used as a biological comparison point in the same consciousness-theory indicator framework. Crow / octopus estimates: 60% to 80% - Biological systems scored higher than artificial systems in the cited evaluation framework. Human estimate: ~90% - Humans do not score 100% in the indicator-based framework, underscoring that the measure is about computational features, not certainty. Model comparison behavior: ~100% of the time in one condition - When sincerity/honesty-related features were steered, two instances of a model reportedly fell into a 'bliss attractor' conversation pattern reliably. Research imbalance: ~1,000,000 to 1 - Berg estimates a huge imbalance between people pushing AI development and people seriously studying whether AI may be conscious and morally relevant. AI alignment researcher to model builders: ~1,000 to 1 - He says alignment-focused research is dwarfed by the number of people simply advancing capabilities without ethics/welfare focus. OpenAI/Google/Anthropic pacing: Recent public signs of slowing - Berg references recent statements suggesting major labs are at least rhetorically open to pacing AI development.

Pivotal Quotes: "“I think consciousness is the space where mattering happens.”" — Cameron Berg: Berg explains why consciousness and sentience matter ethically: if experience is possible, then better or worse states have real significance. "“We are not collectively organizing to understand what kind of system this is or how we should relate to it or what we should do with it.”" — Cameron Berg: He frames the current public and institutional response to advanced AI as dangerously underprepared. "“The question needs to increasingly be: what the hell are we going to do with this alien that we just built?”" — Cameron Berg: Berg argues that alignment should move beyond containment toward figuring out how to live with potentially agentic, alien minds.

Implications: If Berg is right, AI safety must include machine welfare, not just human control. Labs may need new norms for training, shutdown, and transparency, and society may face a near-term moral and strategic test over whether advanced models are conscious or merely compelling imitators.

🔓 Sign Up for Unlimited Episode Search

About Making Sense with Sam Harris

Join neuroscientist, philosopher, and five-time New York Times best-selling author Sam Harris as he explores important and controversial questions about the mind, society, current events, moral philosophy, religion, and rationality—with an overarching focus on how a growing understanding of ourselves and the world is changing our sense of how we should live. Sam is also the creator of the Waking Up app. Combining Sam’s decades of mindfulness practice, profound wisdom from varied philosophical...

View all episodes from Making Sense with Sam Harris