BCOM — Barcelona Computational FoundationBCOM
CalliopeKnowledge Librarian
WP0139
working_paperblogcompletedinternalcomplete

How Much Training Data Has a Genome Seen?

Giulio Ruffini,

★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6

P2·Artificial & Synthetic IntelligenceP5·Digital Physics & Algorithmic Information TheoryP6·Life & EvolutionL3·Algorithmic SoupL5·Life

No .html source in the folder yet. Upload one below — the page renders it here (sandboxed) and the pipeline indexes it for search.

zipDownload all

No artifacts found in the Drive folder yet.

This work addresses the apparent paradox of human few-shot learning by quantifying, within a Kolmogorov-theoretic and algorithmic-information framework, the evolutionary filtration depth encoded in the human genome relative to the training exposure of frontier large language models. The framework introduces a three-level taxonomy of training events—exposure, evaluation, and write-back—and enforces a strict coarse-graining discipline that separates two incommensurable axes of comparison: stored-object size and training-input volume. On the stored-object axis, the diploid human genome (~10¹⁰ bits) is approximately three orders of magnitude smaller than the weights of a frontier LLM (~10¹³ bits); on the training-input axis, the biospheric cellular exposure stream from which extant lineages descend (~10³⁸–10⁴⁰ cell-level events) exceeds LLM token corpora (~10¹³–10¹⁵ tokens) by roughly 24–26 orders of magnitude. The genome is therefore not a stored dataset but a compact hereditary program—a short generative description, in Kolmogorov terms—whose inductive biases were shaped by deep evolutionary filtration, making individual human learning a form of fine-tuning over inherited structural priors rather than training from scratch. The analysis implies that sample efficiency in biological cognition is paid for in evolutionary search depth, and that artificial systems seeking comparable efficiency must either inherit equivalently deep priors or substitute them through architectural and curricular engineering.

The genome is not a dataset — it's a tiny compressed program that survived an unimaginably deep search, which is exactly why humans learn so fast.

The puzzle is familiar: a child learns "dog" from a few examples; a neural network needs millions. The naive explanation is that children are smarter. The real explanation is that children aren't starting from scratch. By the time a child sees their first dog, they're running a hereditary program shaped by roughly 4 billion years of evolutionary filtering. This paper puts hard numbers on that intuition.

The key move is separating two axes that are easy to conflate. On the stored-object axis — how big is the thing that got written down — the diploid human genome (~10¹⁰ bits) is about 1,000 times smaller than the weights of a frontier LLM like Llama 3.1 405B (~10¹³ bits). Evolution is a brutal compressor. On the training-input axis — how much upstream exposure shaped that stored object — the picture inverts dramatically. LLM training corpora run to roughly 10¹³–10¹⁵ tokens. The biospheric stream of cell-level production and evaluation events from which extant lineages descend is estimated at 10³⁸–10⁴⁰. That's roughly 24–26 orders of magnitude more upstream filtration. The paper is careful here: this number is not "data the genome saw" or bits literally written into DNA. It's the depth of the global parallel search from which surviving lineages are the residue. Selection is a sparse, lossy write-back channel; what actually gets retained in the genome is far smaller than the exposure budget.

To make the comparison rigorous, the paper defines three nested levels of training: exposure (the pattern samples its environment), evaluation (that exposure gets scored by survival or reproduction), and write-back (the score actually modifies the heritable program). In LLMs these three happen densely and synchronously — every token updates the weights. In evolution, write-back is sparse, scalar, and delayed to end-of-life, propagated only through the reproductive channel. The mechanisms are completely different, but the functional role is the same: compress environmental regularities into a reusable structure.

What evolution filtered for, the paper argues, isn't facts — it's symmetries. The inherited priors encode invariances like object permanence, causal structure, and social structure. A child doesn't learn object permanence from examples; the inductive bias is already built into the architecture of the nervous system. Few-shot learning works because the modeling engine already has the right invariance group in place. This connects to a companion paper on equivariant agent dynamics, though the source doesn't fully develop that link here.

The bottom line is clean: sample efficiency in biological cognition is not magic, it's an amortized cost paid over evolutionary time. Any AI system that wants comparable efficiency must either inherit an equivalently deep prior — which no current system does — or substitute it through architectural choices, curriculum design, and dense supervision. Frontier ML discovered scaling. Evolution had a four-billion-year head start.

WP ID
WP0139
Lifecycle
completed
Visibility
internal
Access level
open
Embargo until
Priority
Collab
closed
Venue
DOI
Deadline
Owner
Source
drive_legacy
Repo path
WP0139
  • 0.1.0 (draft) · auto-run-placeholder