A Pedagogical Walkthrough of Variational Free Energy
★ Giulio Ruffini, ,
★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6
This note gives a self-contained, pedagogical derivation of variational free energy, starting from the concrete problem of approximating an intractable Bayesian posterior. The central identity is that, for a fixed observation, variational free energy equals the posterior Kullback--Leibler divergence plus surprisal; equivalently, it is an upper bound on negative log evidence, while its negative is the evidence lower bound (ELBO). We explain why this objective is canonical once one fixes a generative model, conditions on data, adopts a KL approximation criterion, and restricts to a tractable variational family --- and we are explicit about the limits of that claim. Later sections turn the abstract expectation into a concrete computation (analytic coordinate ascent, Monte Carlo reparameterization gradients, minibatch stochastic optimization), work a Gaussian example end to end, and connect the same functional form to Helmholtz/Gibbs free energy, the Bogoliubov--Feynman variational bound, path integrals over trajectories, Markov blankets, and active inference. A closing section positions the result within Kolmogorov Theory (KT): variational free energy is the probabilistic, average-case entry point to the same modeling-and-regulation problem that KT treats algorithmically and on single sequences (WP0018, P13, WP0077), and we flag where the two framings agree and where they diverge. WP0176 develops the KT--FEP--AIF bridge in depth; the present note is its pedagogical prerequisite.
Variational free energy is the computable stand-in for an intractable Bayesian objective — here's the full derivation, every connection, and why it matters.
The core problem is simple to state: you have a probabilistic model of how hidden causes produce observations, you see some data, and you want to update your beliefs about those causes. Bayes' theorem gives the exact answer — the posterior — but computing it requires integrating over all possible causes, which is almost always intractable. The standard fix is to pick a tractable family of approximate distributions and find the member closest to the true posterior. The natural closeness measure is KL divergence. The catch: KL divergence to the true posterior still contains the intractable evidence term. Variational free energy is what's left after you notice that this evidence term is just an additive constant in φ — it doesn't affect which approximate posterior wins. So you drop it, and what remains is both computable and has the same minimizer. That's the whole trick.
The resulting objective splits cleanly into two interpretable pieces: an inaccuracy term (how poorly your approximate beliefs predict the data) and a complexity term (how far your beliefs have drifted from the prior, measured as KL divergence). Minimizing free energy means explaining the data well while not over-committing to elaborate beliefs — a built-in Occam's razor. The paper works this through end-to-end for a Gaussian model, where everything becomes closed-form and the solution turns out to be exactly precision-weighted ridge regression. This is a useful anchor: the abstract machinery collapses to something you already know.
The paper then traces the same functional form across several other domains. In statistical mechanics, Helmholtz free energy has the identical structure (expected energy minus entropy), and the Bogoliubov-Feynman variational bound is the same inequality in different notation. For dynamical systems, the latent object becomes a trajectory rather than a point, and free energy becomes expected action minus path entropy — the bridge to Feynman path integrals and to the Free Energy Principle (FEP) for self-organizing systems. The FEP's extra move is to read a physical system's internal dynamics as if they were minimizing free energy over beliefs about external causes — an interpretive claim that requires Markov blanket structure and steady-state assumptions, and which the paper is careful to label "as-if" rather than a logical necessity.
The closing section positions all of this within BCOM's Kolmogorov Theory (KT) program. The complexity term KL[q‖p] is a probabilistic, average-case information cost; Kolmogorov complexity K(x|y) is the length of the shortest program producing x given y. These are related intuitions but not the same thing — statistical dependence between internal and external states does not automatically yield algorithmic shared structure. The paper is explicit: FEP is one probabilistic implementation of the compress-and-predict imperative; the stronger claim that an agent's model shares mutual algorithmic information with the world requires a separate bridge, developed in WP0176. This note is that paper's prerequisite.
- Zenodo
- 10.5281/zenodo.21008520
- WP ID
- WP0028
- Lifecycle
- ongoing
- Visibility
- internal
- Access level
- open
- Embargo until
- —
- Priority
- —
- Collab
- closed
- Venue
- —
- DOI
- —
- Deadline
- —
- Owner
- —
- Source
- drive_legacy
- Repo path
- WP0028 - Free energy tutorial
- v0.3.0 (revision) · cut-version · zenodo:21008521
- v0.2.0 (revision) · cut-version
- v0.1.0 (draft) · drive-legacyAuto-created by Phase 1a bootstrap ingestion.
