Carbohydrates participate in virtually every biological process, yet the chemical diversity of glycans far exceeds that of DNA or proteins, and high-throughput sequencing of these molecules remains out of reach. Mass spectrometry (MS) is the most advanced technique in glycomics, but assigning identities to spectral peaks has long depended on curated experimental databases or manual calculations, which tend to overlook novel glycan compositions. Bligh, et al. (2025) addressed this limitation with GlycoAnnotateR, an open-source R package that performs de novo annotation of glycan compositions directly from MS data. The work was published in the Journal of the American Society for Mass Spectrometry.
Glycan structures are not encoded in the genome but assembled processively by glycosyltransferases and glycosidases, so a glycome cannot be predicted from genomic or proteomic sequence alone. The three principal glycan structure databases—GlyConnect, KEGG GLYCAN, and GlyTouCan—together hold on the order of 250,000 structures, whereas a handful of human glycosyltransferases could theoretically generate more than one million N-glycan structures of 15 monomers or fewer. As a result, comprehensive Glycomics of poorly characterized systems, including marine algal glycans, still relies on laborious manual calculation. A broadly applicable, open-source tool that integrates with existing metabolomics pipelines is therefore needed.
GlycoAnnotateR builds every possible oligosaccharide composition from a set of constraining parameters, then filters the output with a defined set of chemical rules. Its glycoPredict function enumerates monomers and modifications—such as hexose, pentose, sialic acid, sulfate, phosphate, and N-acetyl groups—and returns compositions for oligosaccharides of 1 to 22 monomers in under 10 min using less than 4 GB of RAM. The accompanying glycoAnnotate, glycoMS2Extract, and glycoMS2Annotate functions match experimental and theoretical m/z values and annotate MS/MS fragment ions within R-based workflows. Existing Glycan Profiling platforms benefit from this complementary, untargeted layer of annotation.
Fig. 1 Calculations for the total number of possible compositions for modified hexose and/or pentose-based oligosaccharides. (Bligh, et al. 2025)
In the first use case, GlycoAnnotateR correctly annotated 15 of 16 commercial carbohydrate standards and assigned single, correct structures to five synthetically prepared oligosaccharides in a single 15 ppm step. The tool was then applied to unknown oligosaccharides liberated by enzymatic digestion of Macrocystis pyrifera fucoidan. After filtering adducts, isomers, and in-source fragmentation, the authors identified 18 distinct sulfated mono- and oligosaccharide structures absent from a negative control, demonstrating that dedicated Oligosaccharide Analysis of complex mixtures is now tractable. Because many biologically relevant glycans carry sialic acid, sulfate, or phosphate modifications, each of which can be followed with dedicated Sialic Acid Analysis workflows, the combinatorial approach is well suited to heterogeneous samples.
To test compatibility with imaging mass spectrometry, the authors reanalyzed a published mouse lung MALDI-FTICR data set originally annotated by NGlycDB. GlycoAnnotateR reproduced all 88 NGlycDB annotations and added further candidates with lower mass deviations. When the centroided data were preprocessed with Cardinal, GlycoAnnotateR annotated roughly 15-fold more peaks than NGlycDB, segmenting them into spatial classes. One class, enriched at the tissue edge, contained trisulfated glycans not present in GlyConnect—patterns that warrant confirmation by MS/MS. The study illustrates how mammalian N-Glycan Profiling and spatial Glycomic Characterization can be combined for hypothesis generation in glycoscience.
Fig. 2 GlycoAnnotateR reveals a pattern of ions annotated as trisulfated glycans localized to the edges of a mouse lung tissue section. (Bligh, et al. 2025)
GlycoAnnotateR lowers the barrier to untargeted glycan annotation and complements established R packages for MS data processing, supporting integration of metabolomic and glycomic datasets. By calculating compositions rather than querying a fixed library, it surfaces novel glycan structures that, once validated, can populate the currently sparse glycan databases and feed downstream biomarker discovery.
Reference
About Us
CD BioGlyco is a leading biotechnology company specializing in glycobiology. We deliver high-quality products and services to support cutting-edge research worldwide.