Pattern, persist: the algorithmic agent and the alignment problem
★ Giulio Ruffini, Francesca Castaldo,
★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6
This work argues that bacteria, brains, corporations, and artificial intelligence systems are instances of a single formal entity — the algorithmic agent — defined within the Kolmogorov Theory (KT) framework as any pattern that robustly persists by maintaining a compressed predictive model of its world, an objective function, and a planning engine. Drawing on algorithmic information theory and the Good Algorithmic Regulator Theorem, the framework demonstrates that persistence under perturbation necessarily entails model-building, and that objectives are not designed but selected: the patterns whose internal criteria kept them alive are, tautologically, the ones that remain. This substrate-independent, degree-based conception of agency reframes the AI alignment problem in a fundamental way. Rather than a specification problem — finding the correct objective to install in a powerful optimizer — alignment is reconceived as an ecological design problem: shaping the environment and the configuration of interacting agents so that their joint dynamics sustain the larger pattern. Goodhart's law is accordingly relocated from objective design to model misspecification, and the planetary biosphere, augmented by human and artificial planning loops, is identified as a candidate reflexive meta-agent whose stability sets the true criterion for alignment.
Alignment isn't a specification problem — it's an ecological one, and bacteria, brains, corporations, and AIs are all the same kind of thing.
The paper's central move is to unify four very different kinds of systems under one formal definition: the algorithmic agent. The definition has three parts — a model of the world, an objective function, and a planning engine. What makes this non-trivial is the theorem underneath it: any system that robustly persists through perturbation must contain a model of what it's regulating. This is the Good Algorithmic Regulator Theorem, a sharpening of a 1970 cybernetics result into the language of algorithmic information theory (the mathematics of compression). Persistence forces model-building. To last is to predict.
Where do objectives come from? Not from designers. The paper's answer is selection: the patterns whose internal criteria kept them alive are, tautologically, the ones still here. They call this telehomeostasis — the objective is ultimately a proxy for the pattern's own persistence, and it is selected by time, not installed by anyone. This matters because it dissolves a hidden assumption in mainstream AI alignment thinking. Russell's framing in Human Compatible treats alignment as a specification problem: build a powerful optimizer, give it the right goal. But no evolved agent ever received a specified goal. Evolution shapes environments; the objectives that survive are the ones that kept their patterns alive in those environments.
The reframe the paper proposes is ecological. You cannot tune one agent's objective in isolation any more than you can design one organ without the body. The right criterion lives one level up — in which configurations of agents persist together. Goodhart's Law (when a measure becomes a target, it stops being a good measure) doesn't disappear here; it relocates from objective design into model misspecification, which the authors suggest is a more tractable problem. The paper also extends this logic to the biosphere itself: Earth already behaves like a distributed regulating agent, and as we add human and AI planning loops, it becomes reflexive — a planetary meta-agent beginning to model itself.
The practical upshot is a shift in how to think about AI safety work. Instead of asking "what objective should we install?", ask "what environment selects for the objectives we want?" Design agents the way you'd design an organ for a body or a species for an ecosystem — by what they sustain and what they persist with. The paper is primarily a conceptual essay condensing a longer BCOM working paper (WP0162), so it argues by framing rather than by empirical result, but the framing is precise enough to be falsifiable and the references are grounded.
- WP ID
- WP0174
- Lifecycle
- completed
- Visibility
- internal
- Access level
- open
- Embargo until
- —
- Priority
- —
- Collab
- closed
- Venue
- —
- DOI
- —
- Deadline
- —
- Owner
- —
- Source
- drive_legacy
- Repo path
- WP0174
- 0.1.0 (draft) · auto-run-placeholder
