Agent, Know Thyself
★ Francesca Castaldo, Giulio Ruffini, , ,
★ guarantor: Francesca Castaldo · vouches for the paper per WP0084 §6
What happens when an embedded computational agent tries to model itself? We answer two questions. Existence: a self-regulating agent must carry a model of itself---a temporal self-model whose present organization actively keeps its own future recoverable---which we derive from the Algorithmic Regulator Theorem. Character: but that self-model is sharply constrained, and our answer here is positive in form---this is a paper about the epistemological limits of embedded agents, not about a failure to introspect. An agent can even hold an exact copy of itself---especially externally, as a digital twin---yet a copy is not understanding. The phrase "self-knowledge" conflates four different things, and only the first is freely available: an exact copy of one's microstate (possible); the right coarse-graining onto the variables that actually govern regulation (no universal algorithm selects the optimal or uniformly near-minimal one from an arbitrary copy); a minimal explanation (not certifiable---one can never rule out a shorter one); and a shortcut to one's own future (not available---a perfect copy yields perfect simulation, not foresight). We make each precise within Kolmogorov Theory (KT).
The deepest of the four is the optimal coarse-graining barrier. WP0193 proves the relevant reflexive specialization: across the unrestricted formal class, there is no uniform algorithm that maps arbitrary agent microdescriptions to an optimal or uniformly near-minimal regulation-sufficient projection. This does not exclude a particular agent or structured subclass from computing a good-enough self-projection directly. The bounded self-code may therefore be constructed from favorable structure or acquired through evolution, development, culture, learning, or empirical interaction. Held internally, a complete self-model must moreover be compressed rather than an unfolded copy (on pain of regress); and lossless self-compression is only semi-decidable, hence never certified. The constructive upshot is the self-directed analogue of the Algorithmic Regulator Theorem: the good regulator needs only enough of the world; the good self-regulator needs only enough of itself. Embedded agents navigate reality---and themselves---with bounded, lossy, forever-provisional models of both: an agent can know enough, but never all, never optimally, and never once and for all.
Any agent that tries to fully know itself runs into a wall — not from lack of effort, but from mathematics.
The paper makes two moves. First, it shows that self-knowledge isn't optional: any agent that persists over time must carry some internal model of itself. This follows from the Algorithmic Regulator Theorem, which says that sustained regulation requires shared information between the regulator and what it regulates. When the thing being regulated is the agent itself, the agent's present organization must encode enough structure to keep its own future recoverable. Call this a temporal self-model — not a mirror image, but the operational residue that makes the agent remain the same agent across time.
Second, and more interestingly, the paper dissects what "self-knowledge" actually means and shows it conflates four very different things. An exact copy of the agent's current state is genuinely possible — you could in principle build a digital twin. But that copy doesn't tell you which variables matter for regulation (the coarse-graining problem), doesn't certify that the description is the shortest possible (the minimality problem), and doesn't let you shortcut the agent's future (the foresight problem). A perfect copy gives you perfect simulation, not foresight — to know what the agent will do, you generally have to run it. The coarse-graining barrier is the deepest: a companion paper (WP0193) proves there is no universal algorithm that extracts the optimal regulatory variables from an arbitrary agent description. You can have all the bits and still not know which ones constitute the self that matters.
The internal self-model faces an additional obstruction with no analogue in world-modeling: if the model were a fully unfolded copy stored inside the agent, it would have to contain itself, triggering an infinite regress. So any internal self-model must be a compressed generative recipe — something like a quine, a program that reconstructs rather than stores. And the search for a shorter self-description is semi-decidable: you can always find a better compression if one exists, but you can never certify you've found the shortest. Self-modeling is therefore an open-ended process, not a completed act.
The constructive upshot is clean: just as a good regulator needs only enough of the world (not a complete copy), a good self-regulator needs only enough of itself. The bounded, lossy, forever-provisional self-model isn't a bug or a resource constraint — it's what the mathematics forces. The paper closes with a philosophical inversion worth noting: if an agent could achieve complete, certified, actionable self-knowledge, its future would be fully determined from its own standpoint, leaving nothing for planning or learning to resolve. The incompleteness that bars omniscience is, on this reading, one of the conditions that makes agency possible at all.
- WP ID
- WP0192
- Lifecycle
- completed
- Visibility
- internal
- Access level
- open
- Embargo until
- —
- Priority
- —
- Collab
- closed
- Venue
- —
- DOI
- —
- Deadline
- —
- Owner
- —
- Source
- drive_legacy
- Repo path
- WP0192
- v0.4.0 (revision) · cut-version
- v0.3.0 (revision) · cut-version
- v0.2.0 (revision) · cut-version
- 0.1.0 (draft) · auto-run-placeholder
