University of California, San Diego
Vision-language foundation model linking human brain activation maps and neuroscience text for text-to-brain and brain-to-text generation.
NeuroVLM is a vision-language framework that embeds human neuroimaging data and the neuroscience literature into a shared latent space, enabling movement in both directions between a written description of cognition and the pattern of brain activity it implies. Decades of functional neuroimaging results are reported as coordinate-based activation maps scattered across tens of thousands of separate publications, which makes it difficult to relate an arbitrary text query to the brain regions it engages, or to interpret an unlabeled activation map in the vocabulary of cognitive neuroscience. NeuroVLM treats these two modalities as paired views of the same underlying phenomenon and learns to translate between them.
Unlike retrieval-only approaches, NeuroVLM is generative in both directions: it can synthesize a predicted whole-brain activation map from arbitrary natural language, and it can generate a textual interpretation of an input neuroimage. This places it alongside contrastive language-image models adapted to neuroimaging while extending them with reconstruction objectives that make map synthesis possible. It was developed by researchers at the University of California, San Diego and released as a bioRxiv preprint in early 2026, with an open-source implementation.
The model pairs a fine-tuned SPECTER-based text encoder with a 3D autoencoder that maps coordinate-based activation likelihood volumes into 768-dimensional embeddings, projecting both modalities into a common space trained with an InfoNCE-style contrastive objective alongside reconstruction losses. Training data comprise roughly 30,000 neuroimage-text pairs assembled from large open resources, including approximately 27,000 coordinate-based activation maps and a comparable volume of neuroscience publications sourced from PubMed Central and Neurosynth. Performance was assessed with 10-fold cross-validation, using MSE, SSIM, and Dice to quantify reconstruction fidelity of generated maps and semantic retrieval metrics to evaluate cross-modal matching between text and neuroimages.
NeuroVLM serves cognitive neuroscientists and meta-analysis researchers who need to move between natural language and brain data. It can generate a candidate activation atlas for an arbitrary cognitive description, propose a textual account of an unlabeled map, label functional networks, and power a bidirectional search engine that retrieves publications most relevant to a neuroimage query or images most relevant to a text query. These capabilities support hypothesis generation, literature synthesis, and lesion-to-function reasoning, where a spatial pattern of damage or activity is interpreted in terms of the cognitive functions it implicates.
By unifying neuroimaging and neuroscience text in a single generative model, NeuroVLM offers a way to make the accumulated coordinate-based literature queryable and generative rather than static, connecting written hypotheses to predicted brain activity and back. Its open Apache-2.0 release lowers the barrier for researchers to build on the approach. As a preprint that has not yet completed peer review, its claims await independent validation, and because it is trained on meta-analytic coordinate-based maps rather than raw voxel-level scans, its generated maps reflect the resolution and biases of that aggregated data; the reported evaluations are computational rather than prospective experimental tests.
Hammonds, R., et al. (2026) NeuroVLM: A generative vision-language framework for human neuroimaging. bioRxiv.
DOI: 10.64898/2026.02.06.704508Papers that recently cited this model.
The most-cited papers that cite this model.
Not enough data