Algorithmic Ethics Taxonomy (AB) with Objectives Only
★ Giulio Ruffini,
★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6
This work presents a formal taxonomy of ethical behaviors for algorithmic agents, grounded exclusively in objective functions without appeal to utility theory or external normative constructs. The framework characterizes agent-to-agent interactions along three principal axes: the causal effect of agent A's actions on agent B's objective return (Δ_B), A's awareness of that effect as measured by algorithmic information-theoretic mutual information, and A's targeting or intent as captured by the sensitivity of A's policy to its predictions of B's outcome. Crossing these axes yields a structured classification of behaviors—ranging from circumstantial harm and knowing-but-indifferent harm through instrumental targeted harm and spite, to circumstantial kindness, intentional kindness, and prosocial objective-level coupling—each assigned precise mathematical conditions expressible in terms of episode trajectories and computable from simulation data. The framework further accommodates inequity aversion, reciprocity, CIRL-style paternalistic assistance, autonomy support via empowerment, impact regularization, and Shapley-based credit attribution, all encoded directly within or as penalties on the agent's single objective function. Drawing on causal influence diagrams, Kolmogorov complexity, and behavioral economics, the taxonomy provides drop-in formal definitions suitable for use in methods sections of multi-agent alignment research.
A compact reference taxonomy that classifies every ethically meaningful interaction between two agents using only their reward objectives — no utility theory required.
This is a slide deck, so the content is deliberately terse. What it offers is a precise vocabulary, not a research argument. The core idea is simple: if you have two agents A and B, you can characterize the moral quality of A's behavior toward B by measuring three things — the causal effect of A's actions on B's objective (does B end up better or worse off?), whether A's world-model actually predicts that effect (awareness), and whether A's policy is sensitive to that prediction (targeting/intent). Cross those three dimensions and you get a clean ladder of culpability: from "circumstantial harm" (bad outcome, no awareness, no intent) up through "knowing-but-indifferent" to "targeted instrumental harm" to "spite," where harming B is literally baked into A's own objective function.
The key technical move is the coupling coefficient χ_{A→B}, defined as the partial derivative of A's objective with respect to A's prediction of B's outcome, holding everything else fixed. This single number distinguishes terminal from instrumental social behavior. If χ = 0, any help or harm is a side-effect of A pursuing something else. If χ > 0, A genuinely "wants" B to do well at the objective level — that's altruism. If χ < 0, A is spiteful in a formal sense. This lets you encode the full Fehr-Schmidt / Charness-Rabin family of social preferences (inequity aversion, reciprocity, prosociality) as modifications inside a single objective O_A, without introducing a separate utility concept.
The taxonomy also covers the benefit side symmetrically — circumstantial kindness, intentional instrumental kindness, prosocial objective-level coupling — and adds a handful of richer constructs: CIRL-style paternalistic help (A's objective rises when A better infers B's latent true goal), autonomy support (A's objective rises with B's empowerment, i.e., B's ability to influence its own future), and impact regularization (penalizing A for shrinking the reachable-state space, a standard AI-safety technique). Awareness is measured using AIT mutual information — Kolmogorov-complexity-based rather than Shannon-based — which is a theoretically cleaner but computationally harder choice; the deck notes Shannon MI as an optional substitute.
The practical payoff is the "computable labels" slide: given a dataset of agent trajectories, you can estimate each of these quantities via finite differences, model gradients, and causal intervention comparisons, then automatically assign an ethical label to observed behavior. The source does not make explicit how this taxonomy integrates with any specific training pipeline or audit system, but it is clearly designed as drop-in formal definitions for a methods section — a shared ontology for reasoning about multi-agent ethics without leaving the language of objectives and policies.
- Zenodo
- 10.5281/zenodo.21008571
- WP ID
- WP0053
- Lifecycle
- ongoing
- Visibility
- internal
- Access level
- open
- Embargo until
- —
- Priority
- low
- Collab
- closed
- Venue
- —
- DOI
- —
- Deadline
- —
- Owner
- —
- Source
- drive_legacy
- Repo path
- WP0053 Slides Morality (algorithmic)
- v0.1.0 (draft) · drive-legacy · zenodo:21008572Auto-created by Phase 1a bootstrap ingestion.
