BCOMBCOM
CalliopeKnowledge Librarian
WP0017
working_papercompletedpubliccomplete

Navigating Complexity: How Resource-Limited Agents Derive Probability and Generate Emergence

Giulio Ruffini,

★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6

P2·Artificial & Synthetic IntelligenceP5·Digital Physics & Algorithmic Information TheoryP6·Life & EvolutionL2·MathematicsL3·Algorithmic Soup

In the Kolmogorov Theory (KT) of consciousness, an algorithmic agent is an information-processing system that compresses sensory data into simpler models to plan actions that optimize an objective function, while operating under limited data access, finite computational resources, and the fundamental limits of algorithmic information theory (AIT). We show how these limitations naturally give rise to probability, Bayesian inference, precision, and emergence. Using a toy example of an agent compressing pages from a large library, we recover a weighted multi-model strategy in which probabilistic reasoning and Occam's razor appear as the agent navigates between models. We then introduce precision---the confidence the agent assigns to its model relative to noisy data---as the second-order quantity that arbitrates the trade-off between trusting the prediction and trusting the observation. We formalize precision as inverse-variance weighting of prediction errors at the Comparator and show what it gives the agent: a principled model-updating process carried out by the Updater (a submodule of the Modeling Engine), in which a confidence-dependent gain determines how much each prediction error revises the model --- so that reliable, persistent errors reshape the model while structureless errors are retained as residual noise, and structural learning saturates once the compressible regularity has been captured.

We then connect the picture to Karl Friston's Free Energy Principle and Active Inference, which appear as the variational-Bayesian special case of the bounded-agent story, and flag the main differences rather than collapsing the two. Finally, we propose a formal, agent-centric definition of emergence in terms of coarse-graining and Kolmogorov complexity, and connect it to cellular automata, the renormalization group, and partial models. The result is a unified account in which probability, precision, and emergence are all consequences of an agent's drive to compress and model a noisy world under bounded resources.

Probability, precision, and emergence aren't fundamental axioms — they're what you get when a compression-driven agent hits resource limits.

The paper's central move is to ask: what happens when an agent tries to compress a noisy, diverse world but can't compute the theoretically optimal model? The answer, it turns out, is Bayesian inference. The authors use a vivid toy example: a robot working through a vast library, page by page, with only partial access to each page before it must pick a model. No single model handles novels, research papers, and Latin manuscripts equally well. The rational response is to maintain a weighted ensemble of models — and once you add the Solomonoff prior (shorter programs are exponentially more probable, formalizing Occam's razor), you've derived Bayesian inference from first principles rather than assumed it.

Precision is the paper's second contribution and the most operationally concrete. When a model's prediction disagrees with incoming data, the agent faces a fork: is this mismatch signal (update the model) or noise (ignore it)? Precision — defined as inverse variance, i.e., confidence — resolves this via a gain factor that interpolates between "trust the data" and "trust the model." The math is the scalar Kalman gain, which the authors connect explicitly to predictive coding. The key insight is that this same gain operates at two timescales: fast (updating the current estimate) and slow (revising the model itself). As the model accumulates structure, its precision grows, the gain shrinks, and structural learning saturates — the model has absorbed the compressible regularity and the rest is noise. Attention and gating, in this view, are just precision driven to extremes.

The paper then positions Karl Friston's Free Energy Principle (FEP) as a special case of this broader story rather than a competing framework. FEP's variational free energy — expected prediction error plus a KL-divergence complexity penalty — falls out naturally once you assume the agent approximates a true posterior with a tractable one. Two differences are flagged honestly: KT keeps the objective function separate from the generative model (goals aren't baked into priors), and KT's simplicity drive is algorithmic (Kolmogorov complexity) rather than prior-relative.

Finally, emergence gets a formal agent-centric definition: it occurs when a coarse-graining of microscopically incompressible data yields a macroscopic description with dramatically lower apparent complexity, non-trivial entropy, and — crucially — improved utility for the agent. This isn't just philosophical; it connects to cellular automata (Rule 110 looks random microscopically but has compressible macro-structure), the renormalization group in physics (integrating out microscopic degrees of freedom to find scale-invariant behavior), and Kolmogorov's structure function. The punchline is that emergence is partly in the eye of the beholder: it depends on which coarse-graining the agent applies.

Zenodo
10.5281/zenodo.21008492
Preprint
https://osf.io/preprints/psyarxiv/3xy5d_v3
WP ID
WP0017
Lifecycle
completed
Visibility
public
Access level
open
Embargo until
Priority
Collab
closed
Venue
DOI
Deadline
Owner
Source
drive_legacy
Repo path
WP0017 - From Algorithmic Agent to Probability and Emergence
  • v0.3.1 (revision) · cut-version · zenodo:22081194
    v0.3.1 (title page "August 2026 (v3.1)"): two corrections from Kaiti. Eq. (1) gloss fixed — K(M_i) is the length of the shortest program producing the model M_i itself, not "the shortest program capable of generating the data" (that is K(D), or conditionally K(D|M_i)). §7 chain reworded — noise does not "turn code length into negative log-likelihood": once residual uncertainty is represented probabilistically, the conditional code length is −log p(x|z), noise fixing the likelihood's shape and precision its weight. No other content changes.
  • v0.3.0 (revision) · cut-version · zenodo:21008493
    fixing title etc
  • v0.2.0 (revision) · cut-version
    added Precision/Attention concepts and connection with FEP/AIF
  • v0.1.0 (draft) · drive-legacy
    Auto-created by Phase 1a bootstrap ingestion.