Quantifying Structured Experience: Toward a Computational Phenomenology of Subjective Reports
★ Giulio Ruffini,
★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6
A computational theory of phenomenology requires a way to quantify subjective reports. The last four years have produced a small but coherent body of work that treats first-person language as coordinates in a high-dimensional semantic manifold and uses large language models (LLMs), topic modelling and topological data analysis (TDA) to extract pharmacological, clinical and phenomenological signatures from naturalistic text.\ In this Working Paper we (i)~review the state of the art across three empirical domains --- classical psychedelics, affective and psychotic disorders, and contemplative / pure-awareness practice ---, (ii)~identify the epistemological crisis exposed by the finding that current LLMs can generate "mystical experience" reports indistinguishable from human ones on standard instruments, and (iii)~propose a Kolmogorov-Theory (KT) framework for the principled quantification of structured experience.\ In the KT framework a subjective report is the agent's narrated model-over-model, and its informational content is naturally organised along the three dimensions of structured experience,---,structure (compressibility), breadth (coverage of the input--output stream) and realism (model--data match),---,together with the emotion tuple that maps one-to-one onto the Modeling, Objective and Planning engines. We show how longitudinal persistence of a patient's phenomenological identity can be operationalised as mutual algorithmic information across reports, and how the Lie equivariance of the generative model predicts which directions in semantic latent space are phenomenologically meaningful. We close with a concrete protocol sketch for the Enakd EEG-art-therapy trial and flag three novel risks --- the ontological drift of LLM-simulated qualia, demographic bias in latent-space diagnostics, and the seduction of "mystical-score saturation",---,that any computational phenomenology must face.
How do you measure something you can only ever hear described in words?
That's the puzzle this paper sits inside. Psychedelic trips, depressive episodes, psychotic breaks, moments of pure meditative awareness — all of these are known to us almost entirely through first-person reports. Clinicians have long converted those reports into fixed-category questionnaires (mystical experience scales, depression inventories), but that process throws away most of the texture of what people actually say. The paper's real subject is a fast-moving, four-year-old body of work that instead treats the raw language itself as data: feed reports into a large language model, look at where they land in the model's high-dimensional space of meanings, and see what structure falls out — clusters that track drug pharmacology, topological "loops" that track depressive rumination, decay-and-recovery trajectories that predict clinical improvement. It's a genuinely new instrument, built by pointing existing AI machinery at old philosophical problems.
But there's a landmine buried in this approach, and the paper puts it front and center: if you literally ask an LLM to write a "psychedelic trip report," it can now produce text that scores as more "mystical" on standard scales than most real human accounts. That's damning. It means these instruments — old questionnaires and new latent-space methods alike — may be measuring the shape of language, not the presence of any actual experience behind it. A system with zero inner life can ace the exam. This is what the authors call the epistemological crisis: the measuring stick can't tell reality from a good impression of reality.
Their proposed fix comes from BCOM's Kolmogorov Theory (KT) framework, which models a conscious agent as something that builds compressed predictive models of its world and acts to optimize a scalar sense of value. In this view, a verbal report isn't a direct printout of experience — it's a further compression, a story the agent tells about its own internal model. That story can be scored along three axes: how compressible it is (structure), how much of the agent's experience it covers (breadth), and how well it actually matches independent, non-linguistic evidence like EEG signals or receptor biology (realism). This last axis is the paper's answer to the LLM-mimicry problem: text alone can be faked, but text that has to line up with a simultaneous brainwave signature is much harder to fake. They also sketch how to track a person's psychological continuity over time as "shared algorithmic information" between successive reports, predicting that real clinical turning points (like remission from depression) should show up as sudden breaks in that continuity rather than slow drift.
The paper closes by wiring all this into a real upcoming trial — an EEG-guided art-therapy study for adolescent depression — with specific, falsifiable predictions about how brain-signal complexity, report language, and emotional tone should move together. It's honest about what it hasn't done: this is a literature synthesis plus a theoretical proposal, not a finished experiment, and several of its own open questions (how to normalize these measures across languages and patients, whether the framework generalizes unsupervised) are left unresolved.
- WP ID
- WP0085
- Lifecycle
- ongoing
- Visibility
- internal
- Access level
- open
- Embargo until
- —
- Priority
- —
- Collab
- closed
- Venue
- —
- DOI
- —
- Deadline
- —
- Owner
- —
- Source
- drive_legacy
- Repo path
- WP0085
- v0.2.0 (revision) · cut-versionadded a new paper on food!
- v0.1.0 (draft) · drive-legacyAuto-created on first summarize after upload.
