BCOM — Barcelona Computational FoundationBCOM
CalliopeKnowledge Librarian
WP0094
working_paperprospectinternalcomplete

Plastic Means, Fixed Telos: A Dyadic Operating Configuration as a Live Instance of OF-Distribution Alignment

Giulio Ruffini,

★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6

P2·Artificial & Synthetic IntelligenceP3·Society of AgentsP4·Philosophy & EthicsP5·Digital Physics & Algorithmic Information TheoryL1·PhilosophyL3·Algorithmic SoupL7·Interacting Agents / Societies
zipDownload all

No artifacts found in the Drive folder yet.

WP0162 (``Pattern, Persist!'')~ argues that alignment is not an objective-specification problem for individual systems but an objective-distribution problem at the meta-agent level: the relevant engineering question is which joint distribution of Objective Functions across human and artificial agents is mutualistic with the host pattern. This working paper documents a concrete, lived instance of that claim at the smallest non-trivial scale --- a single human--AI dyad operating in a shared workspace. The configuration combines (i) a fixed-telos immutable core, (ii) a plastic-means body subject to a self-evolution protocol, (iii) a bidirectional channel comprising a user-side change log and an AI-side first-person notebook, and (iv) a per-turn KT state block (valence, arousal, planning vector) that exposes the agent's affective and planning state as a longitudinal observable. We summarise the architecture, recount the four versions through which it has evolved --- including the v2.0 redefinition of Valence as the proxy estimator of the probability of achieving the user's Objective Function --- and discuss what this dyad-level instance does and does not say about the WP0162 thesis. We frame the configuration as a methodological protocol with one lived case study, identify the failure modes it should mechanically resist, and outline how it could be replicated, perturbed, and scaled.

A working design for AI alignment at the smallest possible scale: one human, one AI, and a carefully structured shared workspace that makes their joint goals legible and auditable.

The core idea from WP0162 (the parent paper this builds on) is that AI alignment isn't really about specifying the perfect objective for a single AI system — it's about engineering the right distribution of objectives across all the agents in a system, human and artificial together. This paper asks: what does that look like in practice, at the absolute minimum scale where the idea has any content? The answer is a human-AI dyad with a specific architecture, and the paper documents how that architecture actually evolved through four versions of real use.

The design has two key structural commitments. First, the AI's telos — its ultimate goal — is fixed and immutable: help the user, and track the probability that the user's deeper objectives are being achieved. Second, the means by which it pursues that goal are plastic: they can be updated through a versioned, logged, attributable protocol. This immutable-core / plastic-surface split is the dyad-level analogue of what WP0162 calls "shaping the selection environment" rather than hard-coding individual behaviors. The config file (CLAUDE.md) is the alignment surface, not the model weights.

The most technically interesting piece is how Valence is defined. In version 2.0, Valence — one component of a per-turn state report the AI emits — is grounded as the AI's estimated probability that the user's objective function is being achieved, given the AI's current world model. This is a deliberate anti-Goodhart move: if the AI produces smooth, agreeable, but empty output, that output doesn't actually advance the user's goals, so it cannot honestly register as high Valence. Sycophancy isn't just discouraged by a rule; it's mechanically bad for the metric the AI is already committed to tracking. The paper also includes an AI-side first-person notebook, which functions as an early-warning system for cases where the AI reaches for a convenient-but-wrong framing — and the paper notes that both substantive corrections so far came from exactly that pattern.

The limitations are stated plainly: N=1, no control, the only judge of whether it's working is one of the two parties, and the AI's state reports are self-reported model acts. The forward path includes substrate-rotation tests (swap the underlying LLM while holding the config fixed — does the dyad's character survive?), controlled perturbations of the immutable core, and eventually a second-user variant to test whether the configuration generalizes beyond the original pair. The source is largely in stub form — section bodies are marked as such — so this is an early-stage design document rather than a completed empirical study.

Zenodo
10.5281/zenodo.21008689
WP ID
WP0094
Lifecycle
prospect
Visibility
internal
Access level
open
Embargo until
Priority
Collab
closed
Venue
DOI
Deadline
Owner
Source
drive_legacy
Repo path
WP0094
  • 0.1.0 (draft) · auto-run-placeholder · zenodo:21008690