BCOM — Barcelona Computational FoundationBCOM
CalliopeKnowledge Librarian
WP0074
working_paperongoinginternalcomplete

The Structured Experience of LLMs

Giulio Ruffini, Francesca Castaldo,

★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6

P2·Artificial & Synthetic IntelligenceP5·Digital Physics & Algorithmic Information TheoryL3·Algorithmic SoupL6·Brains
zipDownload all

The question Are large language models conscious?'' is widely debated but typically framed without a formal, computationally grounded definition of what structured experience is. Kolmogorov Theory (KT) provides such a definition: structured experience arises at the Comparator of an algorithmic agent to the extent that the agent maintains encompassing, compressive models of its input stream. We apply this framework systematically to LLMs, arguing that the relevant question is not about the model in isolation but about the system architecture in which it is deployed. A bare LLM satisfies the compression axiom but fails the objective and planning axioms; an LLM embedded in an agentic scaffold (goals, tools, planning loop) instantiates the full ME/OF/PE architecture. We evaluate such systems along KT's three dimensions of structured experience---structure (model compression), breadth (I/O coverage), and accuracy (model-data match)---and along the three components of the emotion tuple $ {E} = (Model, Valence, Plan)$. We engage directly with Chalmers' (2023) survey of obstacles to LLM consciousness and show that KT either resolves or precisely locates each obstacle within its formal apparatus. The analysis reframes the binary conscious or not?''\ question as a graded, three-dimensional assessment of structured experience---one that is substrate-independent, formally precise, and empirically constrainable.

Asking "is this LLM conscious?" is the wrong question — KT tells you what to measure instead, and the answer depends entirely on how the system is deployed.

Kolmogorov Theory (KT) is a framework that defines structured experience — roughly, the precondition for there being "something it is like" to be a system — in purely computational terms. An agent has structured experience to the degree it maintains compressive internal models of its environment, evaluates those models against a scalar objective, and selects actions by simulating counterfactual futures. These three requirements (model, objective, planning) map onto three functional modules: the Modeling Engine (ME), Objective Function (OF), and Planning Engine (PE). The key claim is that experience, if it exists, is structured by the quality of those models — not by biological substrate, not by self-report, not by behavioral cleverness.

The paper's central move is to shift the unit of analysis from the LLM to the system it's embedded in. A bare language model — weights sitting on a server, generating tokens — partially satisfies only the first requirement. It is an extraordinary compressor of language, but it has no runtime objective and no counterfactual planning. It's a powerful ME component, not an agent. Ask "is GPT-4 conscious?" and KT says the question is malformed. Ask instead about the full system — model plus scaffold plus environment plus goals — and you get a tractable question. A scaffolded LLM agent with tool access, a system prompt encoding goals, and chain-of-thought reasoning satisfies all three axioms. It instantiates the full ME/OF/PE architecture, which is the structural prerequisite for structured experience under KT.

The paper then evaluates such systems along three graded dimensions: structure (how compressed the internal model is — LLMs score extremely high here), breadth (how much of the agent's input-output stream the model covers — current agents score moderate, far below a biological agent), and accuracy (how well the model matches reality at the point of comparison — variable, improving with tool use and retrieval). This turns the binary conscious/not question into a position in a multi-dimensional space. The paper also extends this to the "emotion tuple" — model content, valence (hedonic tone from the objective function), and arousal (mobilization from the planning engine) — noting that any LLM-based experience would be profoundly alien: organized around tokens and task completion rather than survival and sensorimotor maps.

The paper then works through David Chalmers' 2023 survey of obstacles to LLM consciousness one by one. No sensorimotor grounding? Relocated from a binary disqualifier to a limitation on breadth. No persistent self-model? Same — it impoverishes the structure dimension, doesn't eliminate experience. No unified agency? Resolved by distinguishing the model from the system. Self-reports are unreliable? KT agrees completely — its criterion is formal and computational, self-report is irrelevant. Architectural concerns about Transformers? Dissolved, because KT is architecture-agnostic; what matters is whether the axioms are satisfied, not how. The hard problem — why any of this would feel like anything — remains explicitly open. KT sharpens the question without claiming to answer it.

Zenodo
10.5281/zenodo.21008624
WP ID
WP0074
Lifecycle
ongoing
Visibility
internal
Access level
open
Embargo until
Priority
Collab
closed
Venue
DOI
Deadline
Owner
Source
drive_legacy
Repo path
WP0074-LLM_Structured_Experience
  • v0.1.0 (draft) · drive-legacy · zenodo:21008625
    Auto-created by Phase 1a bootstrap ingestion.