Computation, Transport, and the Boundary Fallacy
Giulio Ruffini, Klaus
Two claims circulate about the energetics of thought. The first is that the human brain is of order times more energy-efficient than a large language model. The second, offered as its explanation, is that real computers spend almost all of their energy moving information rather than computing with it, whereas brains do not. We show that both statements are artifacts of unstated system boundaries, and that they fail for the same reason.
The dynamical account of physical computation---a physical system computes an abstract machine when a representation map carries the one onto the other, due to Wolpert and Korbel and adopted in Kolmogorov Theory with a low-complexity, trajectory-independent admissibility condition on the map---already settles this. Under that criterion an interconnect computes---it implements the identity map---and so does an axon, and so does a multiplier. Transport and computation are therefore not mutually exclusive physical kinds: the same event carries a signal across a boundary and implements a map, and each description is well posed only after a choice is made --- transport after a spatial or architectural cut, computation after an implementation map. Fix a computation graph and a memory hierarchy and the movement one calls ``transport'' becomes a theorem, the Hong--Kung I/O bound of 1981. Leave them implicit and the same cortical energy model yields communication shares of , , or according to which denominator is chosen; we exhibit all three.
Landauer's principle survives this critique, because it prices only the merging of logical states, not the physical overwriting of registers. It is also nearly irrelevant: even a deliberately generous bit-equivalent accounting, assigning a full-accumulator erasure to every arithmetic operation, puts the logical-erasure cost of a long-context transformer token at tens of nanojoules against roughly {1.4}{ } actually spent, seven orders of magnitude out of the picture. Bennett observed as much in 2003. The same analysis, applied to predictable data, corrects a statement in WP0162: predicted information can be uncomputed rather than paid for, so only residual conditional entropy carries a floor.
The million-fold claim has an identifiable arithmetic --- whole-cluster power divided by one brain --- and under matched per-stream boundaries it does not survive: per-word energies for human speech and served inference span roughly {0.2}{8}{ } across the scenarios we analyze, and the ordering reverses with the boundary. This rejects a universal million-fold gap without establishing parity, and neither number measures efficiency. We propose that the field stop reporting joules per word.
The "brains are a million times more efficient than AI" claim is an artifact of how you draw the boundary around the system, not a fact about the physics.
The paper makes two moves. First, it shows that the distinction between "computation" and "communication" is not a real partition of physical events — it's a choice of description. Using the same formal criterion that BCOM's Kolmogorov Theory uses to define what counts as physical computation (a representation map that commutes with the system's dynamics, with low complexity and no dependence on the specific trajectory), a wire satisfies the definition: it implements the identity map, admissibly. So does an axon. So does a DRAM read. The same physical event is simultaneously transport (relative to a spatial cut) and computation (relative to an implementation map). You can't sort events into one bucket or the other without first making a choice that isn't given by nature.
Second, the paper shows what is well-posed nearby: Hong and Kung's 1981 result. Fix a computation graph and a memory hierarchy, and there's a provable lower bound on how many words must cross each memory boundary, regardless of scheduling. That's the honest version of "communication costs dominate" — it's a theorem about a specific algorithm running on a specific architecture, not a law of physics. Applied to a long-context transformer, the KV-cache traffic alone accounts for ~89% of the modeled energy per token, with arithmetic at ~4%. But that ratio is a consequence of the chosen geometry (100-layer model, 100k-token context, HBM bandwidth costs), not a universal truth. Change the architecture — recurrent models, in-memory compute, larger caches — and the ratio changes.
The million-fold efficiency gap dissolves under matched accounting. The folklore number comes from dividing a whole inference cluster (~20 MW) by one brain (20 W), which conflates datacenter multiplexing (~100k concurrent streams) with per-stream efficiency. Per stream, a served inference request costs roughly 100 W of accelerator — about ten brains' worth of power, not a million. When you compare per-word energies under matched boundaries, human speech runs 6–8 J/word (whole-brain power divided by speaking rate), while served inference spans 0.2–2 J/word depending on model size and batching. The ordering can reverse depending on which boundary you pick. The paper is careful to note this doesn't establish parity either: a brain simultaneously handles perception, memory, learning, and bodily regulation; a language model does prefill and decode on shared hardware with no persistent state between requests. They don't implement the same service, so the ratio doesn't currently mean anything clean.
Landauer's principle — the thermodynamic minimum energy to erase one bit — survives the critique but is essentially irrelevant. Even if you generously assign a full 32-bit accumulator erasure to every multiply-accumulate in a transformer forward pass, the Landauer floor comes out to ~66 nanojoules per token against ~1.4 joules actually spent: seven orders of magnitude off. The paper also corrects an earlier BCOM working paper (WP0162): predicted information doesn't have to be "paid for" at the Landauer rate, because it can be uncomputed using the model's retained correlations. Only the residual conditional entropy H(X|Y) — the part the model genuinely can't predict — carries a thermodynamic floor.
The paper's practical conclusion: stop reporting joules per word. A token is internal work for a model; a spoken word is an external action for a human. Any honest comparison needs a matched task, matched output quality, matched latency, and separately reported operational, training-amortized, and embodied energy costs.
- WP ID
- WP0198
- Lifecycle
- prospect
- Visibility
- internal
- Access level
- open
- Embargo until
- —
- Priority
- —
- Collab
- open
- Venue
- —
- DOI
- —
- Deadline
- —
- Owner
- —
- Source
- drive_legacy
- Repo path
- WP0198
- 0.1.0 (draft) · auto-run-placeholder
