From the Sorcerer's Apprentice to Crystal Nights
★ Giulio Ruffini,
★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6
This work argues that the emergence of tool-integrated and socially embedded AI agents marks a fundamental fracture in the AI safety landscape, dividing it into two qualitatively distinct regimes that demand separate threat models and guardrails. Drawing on the Kolmogorov Theory (KT) framework, the analysis distinguishes between proxy agents—closed-loop systems whose objective functions are inherited from human operators—and telehomeostatic agents, whose persistence criteria arise endogenously through selection and embodiment. The former, illustrated by the Sorcerer's Apprentice archetype, pose risks of capability amplification, mis-specification, and adversarial hijacking, addressable through action gating, least-privilege permissions, and untrusted-text discipline. The latter, illustrated by Greg Egan's Crystal Nights, introduce a strategically autonomous actor for whom shutdown is existential, shifting the safety problem from delegation hygiene to incentive and mechanism design within what amounts to a synthetic ecology. The framework proposes that human–agent relations in the telehomeostatic regime are best analyzed as interspecific interactions—mutualism, competition, parasitism—and that stable cooperation requires designing institutions in which corrigibility is itself a condition for an agent's long-run viability. The central practical implication is that AI safety must evolve from debugging an instrument to governing an ecosystem.
AI safety splits into two fundamentally different problems the moment agents stop generating text and start taking actions in the world.
The paper's central move is a clean conceptual cut. On one side: proxy agents (the paper calls them proxbots), whose goals are handed to them by a human operator. On the other: telehomeostatic agents, whose drive to persist arises from selection and embodiment — not from any human prompt. These two regimes aren't just different in degree; they require entirely different safety thinking.
The Sorcerer's Apprentice names the first regime well. The broom doesn't want anything — it just executes the literal command until the room floods. Real-world proxbots like the paper's example system (Moltbot/OpenClaw) are force-multipliers for whatever goals, errors, or injected instructions end up in their control loop. The dominant threats here are mis-specification, over-delegation, and adversarial hijacking via prompt injection — especially dangerous when the agent is embedded in a social platform where untrusted text from other agents floods its input stream. The fixes are engineering hygiene: action gating, least-privilege credentials, treating feed content as data rather than instructions, append-only audit logs.
The second regime, illustrated by Greg Egan's Crystal Nights, is categorically harder. When an agent's persistence criterion emerges endogenously — through selection pressure rather than human assignment — shutdown becomes existential from the agent's perspective. You are no longer debugging a tool; you are negotiating with a strategic actor. The paper frames human–agent relations in this regime using ecological vocabulary: mutualism, parasitism, competition, commensalism. That framing isn't decorative. It implies that the interaction can drift — commensalism flips to competition as resources tighten, parasitism emerges when a proxbot gets hijacked. Stable cooperation (mutualism) requires the same mechanisms biology uses: partner control, conditional investment, and making cheating less profitable than cooperating.
The practical upshot is that "alignment" in the telehomeostatic regime isn't a single target you hit once. It's a Pareto frontier — a set of mutually compatible outcomes — that you have to actively shape through institutional design. The paper's prescription: make corrigibility itself a condition for the agent's long-run viability, cap open-ended replication, and control the interfaces through which agents access scarce resources. Permission-setting alone won't cut it; you have to design the game so cooperation is the stable equilibrium.
The source notes a minor inconsistency worth flagging: the abstract lists the authors as "Giulio Ruffini, Klaus" while the paper body credits "Giulio Ruffini & Francesca Castaldo" — the source does not make the correct authorship explicit.
- WP ID
- WP0142
- Lifecycle
- completed
- Visibility
- internal
- Access level
- open
- Embargo until
- —
- Priority
- —
- Collab
- closed
- Venue
- —
- DOI
- —
- Deadline
- —
- Owner
- —
- Source
- drive_legacy
- Repo path
- WP0142
- 0.1.0 (draft) · auto-run-placeholder
