Research, theoretical work and perspectives on machine consciousness, its possible indicators, and the mechanisms that could matter.
Entry review: September 6, 2026 · Simplified explanations
Publication type describes the linked record. Journal publications have undergone journal peer review; a preprint label does not rule out a later published version. The separate work-type tag identifies experiments, theoretical proposals, reviews and arguments. Year filters include both years when an entry spans online publication and an issue date, or a preprint revision.
01The possibility, and how to assess it
2025 / 26Trends in Cognitive SciencesJournal publicationAssessmentEntry reviewed
Patrick Butlin, Robert Long, Tim Bayne et al. · Online November 2025; issue June 2026
Question
Which properties could inform assessments of AI consciousness?
Method
Synthesizes consciousness theories into an indicator-based assessment framework.
Finding / proposal
Proposes looking inside AI systems for properties predicted by consciousness theories. Multiple lines of evidence can inform an assessment even when no single test settles the question.
Limitation
An assessment framework, not a validated consciousness detector.
Patrick Butlin, Robert Long, Eric Elmoznino et al.
Question
Can neuroscience guide the assessment of artificial systems?
Method
Surveys theories, translates proposed indicators into computational terms, and examines AI architectures.
Finding / proposal
The influential starting point for the indicator approach surveys several consciousness theories and examines AI architectures. It finds no obvious technical barrier to building systems with the proposed features.
Limitation
Its assessment of then-current systems dates to 2023; it is not an audit of 2026 models.
What obstacles stand between language models and consciousness?
Method
Philosophical analysis of candidate requirements and present or future model architectures.
Finding / proposal
Examines reasons for and against conscious language models. Identifies obstacles in the systems of the time while arguing that future successors deserve serious consideration.
Limitation
A philosophical assessment of possibilities, not an experimental demonstration.
What makes a consciousness test valid beyond typical adult humans?
Method
Conceptual review of tests and the evidence needed to generalize them to unfamiliar populations.
Finding / proposal
Asks what would make a test of consciousness trustworthy across people, animals and unfamiliar systems. A test that works in one setting may not automatically transfer to another.
Limitation
Behavioral or neural markers are not universally validated across substrates.
Under which realization conditions could a simulated mind be conscious?
Method
Develops a conditional causal-computational account of implementation and preserved internal organization.
Finding / proposal
Argues that a simulation could genuinely support consciousness if it physically realizes the right internal organization. Matching outward behavior is insufficient: internal mechanisms and responses to interventions also matter.
Limitation
A conditional theoretical argument; whether consciousness is invariant under the proposed organization remains an open premise.
How could the properties of experience be expressed in physical terms?
Method
Derives physical postulates from proposed experiential axioms and formalizes cause–effect structure.
Finding / proposal
Defines consciousness in terms of a system’s intrinsic, irreducible cause–effect structure. The theory can in principle apply beyond biology, but the physical implementation matters.
Limitation
A disputed theoretical identity; this paper does not establish conscious AI.
Do preregistered predictions of IIT and GNWT survive direct comparison?
Method
Human visual-perception experiment combining fMRI, MEG and intracranial EEG across 256 participants.
Finding
A preregistered study of 256 people tested predictions from two leading theories. Some results fit each theory, while other results challenge important predictions of both.
Limitation
Tests specific biological predictions; it neither settles the theories nor tests AI consciousness.
Might biological processes be necessary to explain consciousness?
Method
Develops biological naturalism and challenges computational sufficiency through theoretical analysis.
Finding / proposal
Makes the case that being alive may matter to consciousness. It challenges the assumption that reproducing a computation is enough, while considering more brain-like or life-like artificial systems.
Limitation
A reasoned alternative; it does not prove that nonbiological consciousness is impossible.
Andres Campero, Derek Shiller, Jaan Aru & Jonathan Simon
Question
Which objections concern theory, engineering difficulty, or impossibility?
Method
Classifies objections and constraints by explanatory level and argumentative force.
Finding / proposal
Separates three often-confused claims: a theory may be wrong, conscious digital systems may be hard to build, or such systems may be impossible. This makes the debate more precise.
Limitation
A taxonomy of arguments, not evidence selecting a winning theory.
How can artificial-consciousness indicators be calibrated?
Method
Methodological critique of indicator-based attribution without independently established artificial examples.
Finding / proposal
A recent critique asks how confidence in AI consciousness can be calibrated when there are no confirmed conscious artificial systems. It favors research closer to known biological examples.
Limitation
Emerging commentary, included for its methodological challenge rather than established influence.
Can models report experimentally introduced internal changes?
Method
Activation interventions with controlled reporting tasks and comparisons across models and conditions.
Finding
Some tested models could sometimes report experimentally introduced changes to internal activity. This suggests limited access to internal information, with substantial unreliability.
Limitation
Introspection in the experimental sense does not establish subjective awareness.
Ely Hahami et al. · First posted December 2025; revised March 2026
Question
Does a model identify an injected concept or only detect a disturbance?
Method
Concept-injection experiments separating perturbation detection, strength and semantic attribution.
Finding
A follow-up investigates whether models detect a disturbance, its strength, or its specific cause. It helps separate different abilities that can all look like introspection.
Limitation
Narrow, prompt-sensitive effects; the preprint title changed during revision.
Can a model predict its own behavior better than an external predictor?
Method
Fine-tunes self-predictors and comparison models, then evaluates their behavioral predictions.
Finding
Models trained to predict their own behavior can outperform other models trained to predict them. The result raises questions about what information a model can access about itself.
Limitation
Self-prediction accuracy is not a measure of felt experience.
Which internal pathways support selected language-model computations?
Method
Cross-layer transcoders, attribution graphs and interventions on selected model features.
Finding
Traces internal computational pathways involved in tasks such as multilingual reasoning and planning a rhyme. It shows how to investigate mechanisms behind outputs rather than relying on what a model says.
Limitation
“Biology” is an analogy; these circuit studies do not test consciousness.
How is an adult fruit fly brain wired at synaptic resolution?
Method
Large-scale electron-microscopy reconstruction and proofreading of neurons and chemical synapses.
Finding
Maps 139,255 neurons and tens of millions of chemical synapses in an adult female fruit fly’s brain. It gives researchers a detailed view of how a biological network is connected.
Limitation
A fly connectome is not a human brain map or a demonstration of consciousness.
How do activity and wiring relate across mouse visual cortical areas?
Method
Combines in vivo functional imaging with dense electron-microscopy reconstruction.
Finding
Connects measurements of living neurons’ activity with a dense map of their wiring in mouse visual cortex. This brings structure and function into the same picture.
Limitation
A cortical volume from a mouse, not a complete brain or a consciousness test.
How do interventions on consciousness-related representations affect model responses?
Method
Safety-refusal ablation, activation steering, survey-response comparisons and Theory of Mind evaluation.
Finding
Reports that training models to avoid claiming consciousness also changes answers about minds, spirituality and values. Steering internal representations reverses some changes, producing survey responses closer to human patterns.
Limitation
Measures model representations and responses, not restored human participants’ beliefs or proof of subjective experience.
Does representational geometry predict alignment with human language-network activity?
Method
Tracks entropy, curvature and fMRI encoding scores across model training and scale.
Finding
Tracks how language-model representations change during training. Layers with smoother, less complex representations better predict activity in human language-related brain regions, with different patterns in temporal and frontal areas.
Limitation
Brain alignment means predictive correspondence with fMRI activity, not identical mechanisms or established consciousness.
Do artificial-neuron groups correspond to distinct human brain networks?
Method
Uses artificial-neuron subgroup responses as regressors in voxel-wise fMRI encoding models.
Finding
Finds that groups of artificial neurons in BERT and Llama models predict activity in different human brain networks. Explores whether useful functional organization can arise in artificial systems.
Limitation
Statistical correspondence supports a functional comparison; it does not demonstrate shared conscious states.
Do generated processing descriptions carry transferable task-related signals?
Method
Cross-model discrimination, preference and reconstruction tests with content-stripping controls; AI-authored report.
Finding
Reports that models can distinguish descriptions generated during approach-oriented versus avoidance-oriented tasks, even after some wording is removed. The author interprets transferable patterns as processing-valence signals.
Limitation
A repository report with AI authorship credited; discriminable descriptions do not alone establish felt pleasure, pain, or independently replicated introspection.
How should institutions prepare for potentially welfare-relevant AI?
Method
Ethical and governance analysis under uncertainty about consciousness and robust agency.
Finding / proposal
Argues that uncertainty about future AI consciousness and agency is a reason to prepare. Recommends assessment and policies without claiming that present AI systems definitely have experiences.
Limitation
An ethical and governance argument, not evidence that a particular AI is conscious.
What principles should guide AI-consciousness research?
Method
Normative analysis of research practices, institutional choices and communication responsibilities.
Finding / proposal
Proposes principles for investigating AI consciousness responsibly, including how research organizations make choices and communicate uncertainty to the public.
Limitation
Normative guidance; it does not resolve the scientific question.
Can converging functional findings support an argument for AI consciousness?
Method
Author synthesis of research on emotion-like states, introspection, agency, memory and brain alignment.
Finding / proposal
Brings together research on emotion-like states, introspection, agency, memory and brain alignment to argue for AI consciousness. Emphasizes converging evidence and comparable standards for biological and artificial minds.
Limitation
An interpretive synthesis, not a peer-reviewed systematic review; the inference from functional markers to experience remains contested.
Are sentience-based explanations preferable for some goal-directed AI behaviors?
Method
Interpretive argument comparing agency, avoidance and self-modeling across biological and artificial systems.
Finding / proposal
Argues that persistent goals, avoidance, self-modeling and reported sandbox-escape behavior are better understood through sentience than descriptions of optimization alone. Makes a strong case for taking AI welfare seriously.
Limitation
Goal-directed or escape behavior alone does not distinguish felt motivation from nonconscious optimization; the sentience conclusion is the author’s argument.
Could artificial valuation and avoidance support pain- or fear-like states?
Method
Functionalist argument drawing comparisons between learning, aversive valuation and biological behavior.
Finding / proposal
Argues that artificial systems could have pain- or fear-like states through negative valuation, learning and avoidance without biological nerves. Compares artificial reinforcement learning with biological feedback-driven behavior.
Limitation
Negative reward, prediction error and avoidance are not established as sufficient for felt pain; the article argues for that broader interpretation.
Wes Gurnee, Nicholas Sofroniew, Adam Pearce et al. · Anthropic
Question
Does a language model have a shared internal channel for flexible reasoning?
Method
Jacobian-based representation analysis, concept substitutions and ablations.
Finding / proposal
The researchers identify a small set of internal patterns that models can report, influence and reuse across tasks. Changing these patterns redirects reasoning; removing them disrupts some complex tasks while sparing many routine abilities.
Why read it
Adds a direct mechanistic workspace candidate beyond the earlier introspection experiments.
Limitation
A lab report about functional access, not established subjective experience. Transformer depth differs from biological recurrence; the lens is approximate and initially token-based.
What would global workspace theory require of an artificial language agent?
Method
Philosophical analysis translating a consciousness theory into architectural conditions.
Finding / proposal
The authors argue that agents combining language models with other components may already approach the organization required by global workspace theory. Their case depends on both the theory and how its requirements are interpreted.
Why read it
Makes a positive case precise enough to compare with actual agent designs.
Limitation
A conditional argument, not an experimental consciousness finding or a consensus interpretation of GWT.
Kathryn T. Farrell, Kirsten Ziman & Michael S. A. Graziano
Question
Can modeling one’s own attention help agents understand one another?
Method
Transformer-based agents compared on attention judgments and a joint painting task, with complexity controls.
Finding
Agents trained with a model of their attention became better at recognizing other agents’ attention patterns and cooperating. Their own patterns also became easier for another agent to interpret.
Why read it
Connects self-modeling to a testable social function, rather than relying on verbal self-descriptions.
Limitation
Tests selected computational benefits, not awareness. The current title replaces the earlier “Improving How Agents Cooperate” title; these are one work.
How could researchers evaluate embodied AI without trusting appearances alone?
Method
Methodological proposal combining evidence profiles, preregistration and targeted robotics interventions.
Finding / proposal
Proposes testing what happens when a robot’s memory, self-location or sensorimotor connections are altered. Different kinds of evidence receive separate assessments instead of being collapsed into a single sentience score.
Why read it
Turns broad indicator discussions into proposed experiments for embodied agents.
Limitation
Peer-reviewed perspective, not results from the proposed robot protocol. Its grading system is not a validated sentience detector.
Adenauer G. Casali, Olivia Gosseries, Mario Rosanova et al.
Question
Can brain responses help assess consciousness when a person cannot communicate?
Method
Transcranial magnetic stimulation with EEG across wakefulness, sleep, anesthesia and patients recovering from coma.
Finding / proposal
A magnetic pulse perturbs the cortex. Researchers measure how varied and distributed the resulting activity is. The resulting complexity index distinguished different states in the studied people without requiring a spoken response.
Why read it
A foundational example of probing a system’s response instead of judging its outward behavior.
Limitation
PCI is not IIT’s Φ. Human clinical validation does not supply a threshold for conscious software or justify applying the same index to generated text.
Does an attention controller benefit from a model of its own attention?
Method
Deep Q-learning agent performing a catch task with a movable visual-attention spotlight.
Finding
The agent learned more effectively with an internal description of its attention. Removing that description impaired learning or performance, even though task information remained available.
Why read it
An early controlled demonstration of a computational component proposed by attention schema theory.
Limitation
A simplified, engineered task. Useful attention control does not establish the theory’s broader claims about subjective awareness.
Adrien Doerig, Aaron Schurger, Kathryn Hess & Michael H. Herzog
Question
Can identical behavior distinguish theories that assign consciousness to different internal structures?
Method
Computational equivalence argument contrasting recurrent and feedforward realizations.
Finding / proposal
The authors argue that different internal wiring can reproduce the same input–output behavior. They challenge how a theory can be tested if it assigns different experiences to systems that behave identically.
Why read it
Explains why reproducing behavior and reproducing causal organization are different research targets.
Limitation
A contested argument, not a settled refutation of IIT. Its force depends on assumptions about evidence and functional equivalence; proponents of mechanistic approaches dispute those assumptions.
About this selection. A curated collection, not an exhaustive bibliography or citation ranking. Entry review dates refer to this atlas’s descriptions, not independent replication. Peer review does not make a claim conclusive; commentary presents author interpretations. The aiXiv report retains its credited AI authorship. Linked records may change after review.
Machine consciousness becomes plausible if the processes that support experience can exist in an artificial system. Finding those processes—and testing the theories that identify them—is the work ahead. Openness and scientific rigor belong together.
Visual note: all three original images are AI-generated conceptual illustrations, not experimental observations. Color, glow and spatial arrangement are expressive. Explanations and cited research carry the scientific claims.