BCOM — Barcelona Computational FoundationBCOM
CalliopeKnowledge Librarian
WP0124
working_paperongoinginternalcomplete

How Much Training Data Has a Genome Seen?

Giulio Ruffini,

★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6

P2·Artificial & Synthetic IntelligenceP5·Digital Physics & Algorithmic Information TheoryP6·Life & EvolutionL3·Algorithmic SoupL5·Life
zipDownload all
PDFWP0124.pdfThe paper — open to read

Why can a child learn many concepts from a few examples, while artificial neural networks often require enormous training corpora? One answer is that the child is not learning from scratch. The human brain is born into the world with a deeply structured biological prior: a body, nervous system, developmental program, perceptual biases, motivational architecture, and learning machinery shaped by evolutionary history. In this sense, the human brain is massively pre-trained --- not by gradient descent on text, but by billions of years of genome--cell--environment interaction.

The extent of this pretraining depends on the coarse-graining scale at which ``training'' is defined. At the microscopic scale, organisms are physical systems composed of enormous numbers of interacting particles; in the spirit of Lloyd's view of the universe as computation, every biological episode involves vast physical information processing. Here, however, we coarse-grain to the biological pattern of interest: the hereditary cellular system that generates viable organisms and brains. At this level, training occurs when the pattern interacts with the rest of the world and the consequences of that interaction are written back, however sparsely, into the future distribution of hereditary generators.

On this definition, the human genome is not a stored dataset of ancestral experience. Its raw sequence contains only 1010 10^{10} bits, far less than the 1014 10^{14}--101510^{15} bits of text used to pretrain a frontier large language model. But this comparison is misleading: the genome is the compressed survivor of an evolutionary computation. Since early cellular life, hereditary patterns have been filtered through roughly 103810^{38}--104010^{40} cell-level persistence/reproduction events. Even assigning only one effective bit of coarse-grained feedback to each such event yields an evolutionary exposure stream exceeding LLM training corpora by 1023 10^{23}--102610^{26} in bits.

The mechanisms are different. LLMs receive dense differentiable error signals over token streams; evolution receives sparse, delayed, population-level feedback through survival, reproduction, development, competition, and ecological interaction. Nevertheless, both systems accumulate compressed structure from data. In Kolmogorov/AIT terms, the genome is not the data but a short generative program shaped by a much longer world-stream: a survival-weighted prior distilled by billions of years of embodied interaction. This evolutionary prior helps explain why human children can learn rapidly from few examples: their learning system is already highly structured before individual experience begins.

Evolution has run a vastly deeper optimization search than any AI training run — and the human genome is the compressed survivor program that resulted.

The core puzzle is simple: a child learns a new word from one or two examples; a large language model needs billions. The obvious explanation is that the child isn't starting from zero. By the time a human infant encounters their first word, they already carry a nervous system, perceptual biases, motivational drives, and learning machinery that were shaped by an enormously long prior optimization process. This paper asks: how long, exactly, and measured in what units?

The key move is to pick the right unit of measurement. You can't just say "billions of years" — that's a duration, not an event count. The paper fixes on cell-level persistence and reproduction events as the natural analog to a training token. A cell divides, survives, or dies; that outcome is the feedback signal. At this coarse-graining scale (not atoms, not molecules, but cells and lineages), the total number of such evaluation events since early life is roughly 10³⁸ to 10⁴⁰. A frontier LLM like Llama 3.1 trains on about 10¹³ tokens. The gap is 25 to 28 orders of magnitude. Even restricting to mammalian evolution alone — roughly 200 million years, with an effective population size around 60,000 — the lineage-level evaluation count lands squarely in the same ballpark as frontier LLM token counts. Mammals, on their own, have already "seen" as much data as GPT-scale models.

But the paper is careful not to overstate the analogy. The mechanisms are completely different. LLMs get dense, differentiable gradient signals on every token — rich, precise feedback at every step. Evolution gets a single sparse scalar signal per lifetime: did this organism reproduce or not? No gradient, no backprop, just differential survival across a population. The genome is also not a parameter vector in a fixed architecture — it encodes the developmental program that builds the architecture. Selection searches over program structure itself, not just parameter values within a fixed graph. The paper frames this using algorithmic information theory: the genome is not a stored database of ancestral experiences (it's only ~10¹⁰ bits, far smaller than any LLM training corpus). It is instead a compressed survivor program — a short generative description that passed through an enormous filtration process and came out the other side.

The practical upshot is a clean reframing of few-shot learning. Human sample efficiency isn't mysterious cognitive magic. It's the consequence of a learner whose priors were distilled through ~10⁴⁰ evaluation events before individual experience even begins. Individual learning is fine-tuning on top of an extraordinarily deep prior, not training from scratch. The paper closes with a direct challenge to AI: any system that wants comparable sample efficiency must either inherit a prior of comparable depth from somewhere, or compensate with engineering — better architecture, curriculum, scaffolding, and dense supervision.

Zenodo
10.5281/zenodo.21008783
WP ID
WP0124
Lifecycle
ongoing
Visibility
internal
Access level
open
Embargo until
Priority
Collab
closed
Venue
DOI
Deadline
Owner
Source
drive_legacy
Repo path
WP0124
  • 0.1.0 (draft) · auto-run-placeholder · zenodo:21008784