BCOM — Barcelona Computational FoundationBCOM
CalliopeKnowledge Librarian
WP0055
working_paperongoinginternalcomplete

Agents as Programs: Static Wiring and Plastic State Program-Level Invariants and Refactorings---from Thermostats and Atari RL to Active Inference

Giulio Ruffini, Francesca Castaldo, ,

★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6

P1·Computational Neuropsychiatry & NeurophenomenologyP2·Artificial & Synthetic IntelligenceP4·Philosophy & EthicsL1·PhilosophyL6·Brains
zipDownload all
PDFmain.pdfThe paper — open to read

We ask what is truly static versus plastic in an algorithmic agent. To avoid category errors between mathematics and biology, we introduce a four-tier ontology separating (i) an abstract control loop (maintain beliefs, score candidate policies, select, act), (ii) static program/wiring (model class, objective functional form, update operators), (iii) physical implementation, and (iv) plastic runtime state (beliefs, parameters, precisions). Using this ontology as a refactoring lens, we map thermostatic control and Atari deep RL (DQN/MuZero) onto a common ME/OF/PE decomposition, emphasizing that scalar valuation is inherently policy-conditional. We then show how active inference instantiates the same abstract loop with a fixed choice of expected free energy as the policy score, making the exploitation--epistemics structure explicit while allowing its weighting to depend on inferred uncertainty and precision parameters. Finally, for evolved biological systems we propose a viability-based reading: telehomeostasis as persistence of an agent-pattern across time, and we complement this with an algorithmic information-theoretic handle (persistence of a compressible, multi-agent description). We illustrate how persistent low valence can arise from failures at distinct tiers (program, state, or wetware), using depression as a conceptual case study, and we discuss how telehomeostasis can be multiscale when agents are composed of sub-agents that sustain group-level patterns.

A unified vocabulary for what stays fixed and what changes in any agent, from thermostats to brains.

The central puzzle is deceptively simple: when we say an agent has a "goal," what exactly is fixed and what can change? The paper argues that most confusion here comes from mixing up four very different things — the abstract logic of a control loop, the specific mathematical rules wired into it, the physical hardware running those rules, and the runtime variables that get updated as the agent operates. These are not the same thing, and conflating them causes real conceptual damage. The paper's main contribution is a clean four-tier ontology that separates them, then uses it as a lens to dissect thermostats, Atari RL agents (DQN and MuZero), and active inference agents, showing they all share the same abstract skeleton.

The skeleton is this: form a model of the world and yourself under a candidate plan, score that plan with a scalar objective, pick the best plan, act, update. What differs across systems is the mathematical wiring — the specific objective functional, the update rules, the model class. In a thermostat, the wiring is a setpoint and a bang-bang rule; in DQN, it's a return functional and TD-learning; in active inference, it's expected free energy under preference priors. The abstract loop is identical. The paper formalizes this with a "policy-conditional valuation" — the key insight that you cannot assign a scalar value to a situation without committing to a candidate plan, because the future you're evaluating depends on what you're going to do. This is not a quirk of one framework; it shows up identically in action-value functions (RL), valence (control theory), and expected free energy (active inference).

Active inference gets special attention because it makes something elegant explicit: the epistemic drive to explore and reduce uncertainty is not a separately programmed bonus. It falls out mathematically from the single choice to score policies by expected free energy under preference priors. When the agent is uncertain, information gain differentiates policies and drives exploration automatically; when uncertainty is low, preference satisfaction dominates. The exploration-exploitation balance is not a tunable knob in the static wiring — it's a consequence of plastic runtime variables like inferred precisions.

For biological agents, the paper proposes that the static objective is persistence — specifically, the log-probability that the agent-pattern survives across time (telehomeostasis). This is deliberately substrate-agnostic: cells replace themselves, organisms reproduce, but the compressed description of the pattern persists. The paper uses depression as a case study, showing that persistently low valence (the agent's best available plan still looks bad) can originate at any of the four tiers — wrong models, miscalibrated valuation circuitry, impaired planning, or actual hardware damage — and these failure modes interact in self-reinforcing loops. The same ontology extends to collectives: a group can be treated as a higher-level agent whose persistence criterion gets encoded either as a collective objective or as expanded state variables inside individual agents' models.

Zenodo
10.5281/zenodo.21008580
WP ID
WP0055
Lifecycle
ongoing
Visibility
internal
Access level
open
Embargo until
Priority
Collab
closed
Venue
DOI
Deadline
Owner
Source
drive_legacy
Repo path
WP0055 - Agents as Programs: Static Wiring and Plastic State
  • v0.1.0 (draft) · drive-legacy · zenodo:21008581
    Auto-created by Phase 1a bootstrap ingestion.