Fine-tuned Prosit predictor of spectra and retention time for citrullinated peptides, separating them from isobaric deamidation in MS searches.
No providers recorded yet. Browse all providers
Citrullination converts an arginine into citrulline and adds 0.9840 Da. So does deamidation of asparagine or glutamine, and so does a search engine picking the wrong monoisotopic peak. Three different events therefore produce the same precursor mass shift, and no amount of mass accuracy separates them. The usual remedy has been manual inspection of every candidate spectrum, which does not survive contact with a tissue-scale dataset — and the diagnostic ion that would settle it, the isocyanic acid neutral loss, is frequently missing because citrulline suppresses the C-terminal y-ions that would carry it.
Prosit-Cit answers this with the one signal that genuinely differs between the three hypotheses: the fragmentation pattern itself. Citrullinated and deamidated forms of the same sequence produce measurably different b- and y-ion intensity distributions, so a model that predicts a spectrum accurately enough can score the competing interpretations against the observed one. Developed by Mathias Wilhelm's computational mass spectrometry group and Chien-Yun Lee's group at the Technical University of Munich, posted as a preprint in October 2024 and published in Molecular & Cellular Proteomics in 2025, Prosit-Cit is a fine-tuned member of the Prosit spectral-prediction family alongside the catalog's Prosit-PTM and Prosit-XL.
The model is one half of the contribution; the other is the rescoring procedure built around it. For every citrullinated or deamidated match, the pipeline explicitly generates the alternative readings of the same scan — the other +0.98 Da placements plus the unmodified form — and makes the model choose between them.
Prosit_2024_intensity_cit
and Prosit_2024_irt_cit, and wired into the Oktoberfest rescoring package with a published
tutorial notebook.Prosit-Cit keeps the encoder–decoder recurrent architecture of the Prosit lineage, mapping peptide sequence, precursor charge, collision energy, and fragmentation type to b- and y-ion intensities at charges 1–3 for sequences up to 30 residues. Training drew on ProteomeTools reference spectra — over 30 million spectra from more than 500,000 tryptic and 242,000 non-tryptic peptides — extended with roughly 53,000 spectra from about 2,500 synthetic citrullinated peptides of known site. Deamidated asparagine and glutamine are represented as aspartate and glutamate, so no separate deamidated training peptides were required. Held out from training were 205 citrullinated peptides (~14,000 spectra) for validation and 105 peptides (~7,300 spectra) as a holdout set. Median spectral angle reached 0.85 for HCD and 0.78 for CID on citrullinated peptides, against 0.92 and 0.90 for unmodified ones; retention-time prediction gave Pearson R ≥ 0.99 with a Δ-iRT 95% window of 2.25–2.35 minutes for citrullinated peptides. On a dilution series of synthetic citrullinated peptides spiked into a human cell lysate, MSFragger plus Prosit-Cit rescoring held precision above 80% across concentrations where the search engines alone fall below 50%.
Because the model rescores search output rather than changing acquisition, it applies to already-published DDA data. The authors reanalyzed ten human tissue proteomes, recovering most previously manually validated sites while identifying substantially more, and then applied the same pipeline to an Arabidopsis thaliana atlas, finding 199 citrullination sites on 169 proteins across 30 tissues — the first large-scale citrullination map in a plant, with the highest counts in hypocotyl and floral organs. Any group holding proteomics data and an interest in citrullination — in rheumatoid arthritis, neutrophil biology, or plant stress responses — can run the pipeline without new instrument time.
Prosit-Cit shows that a modification-specific spectral predictor plus careful FDR design can turn a PTM previously restricted to hand-curated case studies into a proteome-scale measurement, reporting up to 14 times more citrullination sites than prior analyses of the same data. The honest limits are stated by the authors: no enrichment step precedes measurement, so identified sites skew toward abundant proteins, and a cold-stress flower experiment yielded only six sites despite 12,000-protein depth. Spectral angles on citrullinated peptides remain below those on unmodified ones, a consequence of the comparatively small citrullinated training set. Selective enrichment chemistry, not prediction, is the next bottleneck.
Much of this page is generated or calculated automatically. Flag anything that looks off and we will re-run it.