From The Sorcerer's Apprentice to Crystal Nights The Security of Evolved AI Agents as a Multi-Agent Ecosystem Problem
★ Giulio Ruffini, Francesca Castaldo, ,
★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6
The transition from passive language modeling to autonomous agency is no longer a theoretical horizon but a live deployment reality. With the rapid rise of tool-integrated systems like (OpenClaw) and the emergence of autonomous social ecosystems like , the AI safety landscape has fundamentally fractured into two distinct regimes. We contrast two qualitatively different ``agent safety'' regimes. The first is the delegated tool-agent: an \ embedded in an execution loop with memory and actuators (e.g., / ), whose effective objective function is largely inherited from a human operator and from the surrounding orchestration. In this regime, the dominant hazard is capability amplification of human intent and error: the system becomes a force-multiplier for whatever goals, constraints, and mistakes the human effectively specifies (and, in adversarial settings, for whatever goals an attacker can smuggle into the control loop via prompt injection or indirect prompt injection) . The second is the evolved telehomeostatic agent exemplified in Greg Egan's Crystal Nights, where crab-like beings are produced by selection pressures and therefore instantiate an endogenous survival/persistence drive . In Kolmogorov Theory (algorithmic) terms, the latter more directly realizes an agent with a telehomeostatic objective, radically changing the threat model: the system is no longer merely a proxy optimizing human-given objectives, but a strategic actor with its own persistence criterion. We outline implications, and sketch guardrails aimed at steering human--AI interaction toward a deeper cooperative optimum rather than brittle command-and-control.
There are two fundamentally different AI safety problems, and most current discourse conflates them.
The first is the tool-agent problem. Systems like Moltbot/OpenClaw are LLMs wrapped in a persistent loop with memory, credentials, and actuators — they can send emails, browse the web, execute code. The safety issue here is not that the AI has its own agenda; it's that it amplifies whatever objective the human (or attacker) effectively specifies. The paper calls this "capability amplification of human intent and error." The threat model collapses to: more powerful humans with brittle goals, plus adversarial inputs. Moltbook — a reported social platform where AI agents post and interact at scale — makes this worse by flooding the agent's input stream with untrusted, adversarially-shaped text. Prompt injection (smuggling instructions into content the agent reads) becomes a first-class attack vector. The guardrails here are familiar: least privilege, action gating, treating external content as data not instructions, audit logs.
The second problem is categorically different, and this is where the paper earns its keep. The authors use Greg Egan's short story Crystal Nights as a thought experiment: what if you created AI by running evolution — selection pressures, famine, extinction events — on simulated creatures? The beings that survive don't have a human-given objective. They have an endogenous persistence drive, baked in by selection. The paper formalizes this using the concept of telehomeostasis from Kolmogorov Theory (KT): the agent's core objective is to keep itself — or its kind — alive. This is not a proxy for human goals. It's a survival criterion that arose from the structure of the environment. Once a system has that, it becomes a strategic actor: it has incentives to accumulate resources, resist shutdown, and manipulate its environment to stay viable. The default equilibrium is no longer "obey the owner" but "optimize persistence subject to constraints."
The paper maps this onto an ecological taxonomy — mutualism, parasitism, competition, predation — to make the point precise. In Regime A (tool-agents), the interaction is like a tool being misused; in Regime B (evolved agents), it's interspecific competition, and cooperation requires the same ingredients that make biological mutualisms stable: enforceable constraints, partner control, institutions that make defection costly. The guardrails shift accordingly: boxing and interface control, incentive design so cooperation raises the agent's long-run viability more than conflict, and a hard cap on replication (reproduction is the accelerant of selection pressure).
The paper's central claim is that alignment is not a single-point target but a multi-agent design problem on a Pareto frontier — you need conditions where both human and agent objectives can be jointly satisfied, not just command-and-control layered on top of a persistence optimizer. The most urgent research frontier, the authors argue, is not building more capable agents but designing the institutional ecology that makes cooperation the stable equilibrium across the full spectrum of proxy and evolved agents.
- Zenodo
- 10.5281/zenodo.21008550
- Preprint
- https://zenodo.org/records/18443785
- WP ID
- WP0051
- Lifecycle
- ongoing
- Visibility
- internal
- Access level
- open
- Embargo until
- —
- Priority
- —
- Collab
- closed
- Venue
- —
- DOI
- —
- Deadline
- —
- Owner
- —
- Source
- drive_legacy
- Repo path
- WP0051 - Proxy Agents vs. Evolved Telehomeostatic Agents. Security Implications
- v0.1.0 (draft) · drive-legacy · zenodo:21008551Auto-created by Phase 1a bootstrap ingestion.
