Graph diffusion transformer for in-context molecular design, adapting to new tasks from a few molecule-property demonstrations without fine-tuning.
Multimodal biomedical framework aligning frozen single-cell and protein model encoders to an LLM's embedding space for zero-shot reasoning.
Molecular conformer generation from 2D graphs with a diffusion transformer that replaces equivariant layers with graph shortest-path attention biases.
Open-source framework for building RNA and DNA foundation models, featuring WCED pretraining for transcriptomics and SNP-aware encoding for genomics.
Mila / Université de Montréal / McGill University / IBM Research / HEC Montréal
Released May 30, 2025
Protein conformation ensemble generation aligned to force-field energies, calibrating an AlphaFold 3-style diffusion model against MD thermodynamics.
Imaging-genetics foundation model pairing SNP genotypes with brain-MRI phenotypes by contrastive learning to surface many-to-many associations.
Multi-modal, multi-task biological foundation model trained on 2 billion samples spanning proteins, small molecules, and single-cell gene expression.
Molecular foundation model that late-fuses graph, image, and SMILES encoders into one embedding for molecular property and drug target prediction.
Multimodal LLM for inverse molecular design, interleaving text and graph generation with a diffusion transformer and A* retrosynthetic planning.
IBM Research / University of Bern / Inselspital, Bern University Hospital / Medical University of Vienna / Lindenhofspital Bern / University of Duisburg-Essen / ETH Zurich / Lausanne University Hospital / University of Lausanne
Released December 1, 2023
Generative toolkit that synthesizes six aligned immunohistochemistry markers from one H&E histopathology image, trained on unpaired stains.
Geometric relational graph neural network that encodes 3D protein structures through geometry-aware message passing and self-supervised pretraining.
Antibody CDR design model that reprograms a frozen English BERT for sequence infilling, avoiding training a dedicated protein language model.
Large-scale chemical language model trained on 1.1 billion SMILES strings using linear attention transformers for molecular property prediction.