BCOM — Barcelona Computational FoundationBCOM
CalliopeKnowledge Librarian
WP0192
working_papercompletedinternalcomplete

Agent Know Thyself

Francesca Castaldo, Giulio Ruffini, , ,

★ guarantor: Francesca Castaldo · vouches for the paper per WP0084 §6

📁Internal folder
P1·Computational Neuropsychiatry & NeurophenomenologyP2·Artificial & Synthetic IntelligenceP5·Digital Physics & Algorithmic Information TheoryP6·Life & EvolutionL2·MathematicsL3·Algorithmic SoupL5·LifeL6·Brains
zipDownload all

No artifacts found in the Drive folder yet.

What happens when an embedded computational agent tries to model itself? We answer two questions. Existence: a self-regulating agent must carry a model of itself---a temporal self-model whose present organization actively keeps its own future recoverable---which we derive from the Algorithmic Regulator Theorem. Character: but that self-model is sharply constrained, and our answer here is positive in form---this is a paper about the epistemological limits of embedded agents, not about a failure to introspect. An agent can even hold an exact copy of itself---especially externally, as a digital twin---yet a copy is not understanding. The phrase ``self-knowledge'' conflates four different things, and only the first is freely available: an exact copy of one's microstate (possible); the right coarse-graining onto the variables that actually govern regulation (not algorithmically derivable from the copy); a minimal explanation (not certifiable---one can never rule out a shorter one); and a shortcut to one's own future (not available---a perfect copy yields perfect simulation, not foresight). We make each precise within Kolmogorov Theory (KT).

The deepest of the four is the coarse-graining barrier. By algorithmic emergence (``reduction construction''), even given a perfect micro-copy an agent cannot in general compute which of its variables are its regulatory self: the bounded self-code is algorithmically emergent and must be acquired---through evolution, development, culture, or learning---rather than deduced from the substrate. Held internally, a complete self-model must moreover be compressed rather than an unfolded copy (on pain of regress); and lossless self-compression is only semi-decidable, hence never certified. The constructive upshot is the self-directed analogue of the Algorithmic Regulator Theorem: the good regulator needs only enough of the world; the good self-regulator needs only enough of itself. Embedded agents navigate reality---and themselves---with bounded, lossy, forever-provisional models of both: an agent can know enough, but never all, never optimally, and never once and for all.

You can have a perfect copy of yourself and still not know yourself — that's the core surprise this paper unpacks.

The paper asks what happens when an agent tries to model itself, and splits the question cleanly into two halves. First, existence: does a self-regulating agent need a self-model at all? The answer is yes, derived from the Algorithmic Regulator Theorem — a pattern that actively keeps itself recoverable across time must carry internal structure that makes its own future compressible. Second, character: what kind of self-model can it actually have? This is where things get interesting.

The word "self-knowledge" turns out to bundle four very different things. An exact copy of your current state is genuinely possible — you could, in principle, have a perfect digital twin. But three other things are not available: knowing which of your variables are the ones that actually govern your regulation (the right coarse-graining); certifying that your self-description is the shortest possible (verifiable minimality); and getting a shortcut to your own future without just running yourself forward (foresight). The paper makes each of these precise using Kolmogorov complexity — the mathematical theory of shortest descriptions. The deepest of the four is the coarse-graining barrier. Even with a perfect micro-copy in hand, there is no general algorithm that tells you which variables are your regulatory self. That projection has to be acquired — through evolution, development, learning, culture — not computed from the substrate. This is what the paper calls algorithmic emergence: the macro-level description cannot be derived by reduction from the micro-level one.

There is also a reflexive wrinkle with no analogue in world-modeling. A complete self-model held inside the agent can't be a literal unfolded copy — that would require the copy to contain itself, triggering an infinite regress. So any internal self-model must be a compressed recipe, a kind of lossy self-quine. And because certifying that you've found the shortest such recipe would require computing Kolmogorov complexity (which is uncomputable), the search for a better self-description is open-ended and never stamps itself finished.

The constructive upshot reframes what looks like a limitation. A bounded, lossy, forever-provisional self-model isn't a bug — it's what the math forces, and it's sufficient for persistence and regulation. The paper's closing move is almost philosophical: if an agent did have complete, certified, actionable self-knowledge, its future would be a lookup table and its alternatives no longer genuinely open. The incompleteness that bars omniscience is, on this reading, one of the conditions that makes agency possible at all.

WP ID
WP0192
Lifecycle
completed
Visibility
internal
Access level
open
Embargo until
Priority
Collab
closed
Venue
DOI
Deadline
Owner
Source
drive_legacy
Repo path
WP0192
  • v0.3.0 (revision) · cut-version
  • v0.2.0 (revision) · cut-version
  • 0.1.0 (draft) · auto-run-placeholder