Leveraging EEG and fMRI Foundation Models for MDD
Giulio Ruffini, Francesca Castaldo
We propose augmenting SYNERGIA's existing EEG and fMRI biomarker pipeline with representations extracted from pre-trained foundation models (FMs)---large-scale transformer and graph-neural-network architectures trained via self-supervised learning on massive heterogeneous neuroimaging corpora. The core insight is that representation learning (the expensive part) has already been done by these open-source models; what remains is lightweight supervised fine-tuning that SYNERGIA's sample can support. We survey the rapidly maturing landscape of EEG-FMs (REVE, LaBraM, BrainOmni, GEFM, NeuroLM) and fMRI-FMs (NeuroSTORM, BrainGFM, fMRI-LM), and propose a concrete transfer-learning methodology: model selection, preprocessing adaptation, fine-tuning for MDD biotype classification and treatment-response prediction, and systematic benchmarking against classical hand-crafted features and mechanistic whole-brain models. The result is a three-level biomarker architecture (L1:~classical features, L2:~FM embeddings, L3:~digital-twin parameters) in which convergence across levels provides mutual validation and divergence reveals blind spots. We include a critical assessment of risks---including the gap between pre-training distributions and SYNERGIA's intervention-specific data, the challenge of interpretability, and the ``foundation model hype cycle'' in neuroimaging---and discuss connections to the Kolmogorov Theory (KT) framework for algorithmic agents.
Pre-trained brain models already learned the hard part — SYNERGIA just needs to fine-tune them on 120 patients to find depression biomarkers that classical EEG/fMRI features might miss.
The core problem is a familiar one in clinical neuroscience: you want to predict who will respond to treatment, but you have a small trial (N=120), messy multi-site data, and brain signals that are notoriously high-dimensional. Classical hand-crafted features — things like alpha frequency peaks, spectral slopes, and pairwise connectivity — are well-understood but hit a ceiling. They can't capture nonlinear, cross-frequency interactions, and they're brittle when electrode setups differ across sites. Building a deep learning model from scratch on 120 subjects would overfit immediately.
The proposed solution borrows from the same playbook that made large language models useful for specialized tasks: don't train from scratch, fine-tune. A new generation of "foundation models" for EEG and fMRI — transformers and graph networks pre-trained on tens of thousands of subjects via self-supervised objectives like masked reconstruction — have already learned general-purpose representations of brain dynamics. The expensive part is done. What remains is attaching a small task-specific head and fine-tuning on SYNERGIA's labeled data. The paper surveys the most relevant open-source options: REVE (25,000 subjects, 92 datasets, handles arbitrary electrode layouts — ideal for multi-site) and LaBraM for EEG; BrainGFM (pre-trained on 25 psychiatric disorders including depression) and NeuroSTORM for fMRI.
The architecture proposed is deliberately layered. Level 1 is classical features — interpretable, regulatory-friendly, the existing pipeline. Level 2 is foundation model embeddings — richer, data-driven, potentially capturing what L1 misses. Level 3 is mechanistic whole-brain models (Hopf oscillators, graph effective connectivity) — causally interpretable, theory-grounded. The key scientific bet is that convergence across all three levels is strong evidence; divergence is informative about blind spots. If an FM latent dimension predicts treatment response but doesn't correlate with any mechanistic parameter, that's a signal the mechanistic model is missing something real.
The paper is honest about the risks. Domain shift is the biggest one: these foundation models were pre-trained on resting-state healthy brains, and SYNERGIA's patients are depressed and receiving psilocybin plus brain stimulation — a distribution the models almost certainly never saw. The authors argue that even a poorly calibrated FM can detect genuine change by tracking trajectories in latent space rather than absolute positions. The N=120 bottleneck is real too; the mitigation is freezing most backbone layers and fine-tuning only the task head, reducing trainable parameters from millions to thousands. And the paper explicitly flags the hype cycle: many of these models are 2025–2026 preprints with limited independent replication, and benchmark accuracy on curated depression datasets does not automatically translate to treatment-response prediction in a real trial. The FM pathway is designed as augmentation, not dependency — if it underperforms classical features, the core pipeline is unaffected and the negative result is worth reporting.
- Zenodo
- 10.5281/zenodo.21008616
- WP ID
- WP0071
- Lifecycle
- ongoing
- Visibility
- internal
- Access level
- open
- Embargo until
- —
- Priority
- —
- Collab
- closed
- Venue
- —
- DOI
- —
- Deadline
- —
- Owner
- —
- Source
- drive_legacy
- Repo path
- WP0071-EEG_fMRI_Foundation_Models_MDD
- v0.1.0 (draft) · drive-legacy · zenodo:21008617Auto-created by Phase 1a bootstrap ingestion.
