Mass spectrometry foundation model for untargeted metabolomics and lipidomics, naming and quantifying molecules with no reference library.
No providers recorded yet. Browse all providers
An untargeted LC-MS run produces tens of thousands of features, and the conventional way to name any of them is to match their fragmentation pattern against a reference library. That library only contains what somebody already bought as an authentic standard and ran on an instrument, so everything else comes back unidentified — and the features that are named come back as relative peak areas, not concentrations, so two samples measured on different days cannot be put on the same axis. LSM-3 is built to remove both constraints at once: it reads raw high-resolution LC-MS output and returns the structures and the concentrations of what is in the sample, without needing to have seen a given molecule before. Matterworks calls this de novo chemotyping, by analogy with genotyping a genome no one has sequenced before.
LSM-3 is the current generation of the Large Spectral Model line from Matterworks (Somerville, MA), announced on 28 July 2026 alongside the launch of the company's PyxisLabs service. It follows LSM1-MS2 (ChemRxiv, 2024) and LSM-MS2 (arXiv, 2025), which learned an embedding space over MS/MS spectra for identification and downstream biological readouts. LSM-3 is a generation-later sibling of those models rather than a revision of them: it is described as covering both small molecules and lipids, working from MS1 as well as MS2 acquisitions, and — the capability the earlier generation does not have — determining concentration rather than identity alone.
No preprint, technical report, or model card describes LSM-3; what is known comes from Matterworks' own announcement and product documentation, and from trade coverage reprinting them. Matterworks describes its Pyxis interface as fronting "the collection of Matterworks foundation and task-specific expert models," and whether the name LSM-3 designates one checkpoint or that collection has not been stated publicly.
Matterworks states that the current-generation LSM was trained on more than 10 billion spectra spanning millions of distinct molecular structures and biological contexts. That figure is company self-reported — it appears in the launch announcement and on the company's own Approach page, and trade outlets reproduce it from the release — and has not been independently verified. No architecture, parameter count, spectral tokenization scheme, or training objective has been published for LSM-3, and no benchmark results have been released for it. The quantitative figures circulating for the LSM family — isomer-resolution and identification-rate improvements, and a reference library of roughly 1.8 million spectra — belong to LSM-MS2 and describe that earlier, separately published model, not this one. Matterworks' related patent family is WO2025019764A1. The enterprise offering supplies standardized methods, columns, and universal calibrants so customer-generated data reaches the model in the form it expects.
The intended users are laboratories doing metabolomic, lipidomic, and exposomic characterization: biomarker discovery, quantitative mechanism-of-action work, target discovery, lead identification and optimization, and population studies. Access is entirely commercial, through two routes. Samples can be sent to the PyxisLabs service at $200 per sample, covering acquisition, identification, concentration determination, and interpretation delivered through the Pyxis co-scientist interface. Alternatively, a laboratory running its own high-resolution LC-MS instruments can license the model as a subscription starting around $25 per sample, exclusive of instrument operation. There are no downloadable weights, no source code, no public license document, and no free API.
LSM-3's significance lies in what it attempts rather than in any published result: carrying the pretrain-then-apply pattern into raw instrument output for the biochemical layer, and pushing past identification into quantitation, which is the step that makes untargeted measurements comparable across studies. If it works as described, it addresses a bottleneck that has held metabolomics well short of the scale genomics reached. That assessment stays provisional. This generation has no peer-reviewed or preprint description, no released benchmarks, and no independent evaluation; its headline training-scale number is the company's own; and the model is reachable only through a paid commercial service. Matterworks has published on this line before, so a technical account may follow — until one does, the entry rests on vendor claims and on the terms under which the model is sold.
Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.