banner
Atom-Level AI Predicts Glycan–Lectin Binding

Atom-Level AI Predicts Glycan–Lectin Binding

July 17, 2026

Glycans coat nearly every living cell and act as information-rich barcodes that proteins across the tree of life read to distinguish friend from foe. Predicting which glycan a given glycan-binding protein will recognize has long depended on painstaking structural biology. A recent Science Advances study reframes the problem with a lightweight machine-learning model named MCNet, which learns to predict protein–glycan interactions directly from the atomic makeup of the sugar (Carpenter et al., 2025). The work is notable not only for its accuracy but for what it can forecast beyond the training data: how rare, mirror-image sugars would behave at the binding site.

Why Glycan Recognition Is Hard to Model

Unlike DNA, RNA, or proteins, glycans assemble as linear and branched chains whose monomers differ in stereochemistry and linkage, producing a feasible complexity that dwarfs that of comparably sized biopolymers. The human glycome is nonetheless estimated at roughly 7,000 structures, and most experimental data, including glycan microarrays such as those produced by the Consortium for Functional Glycomics, skew heavily toward biologically common, D-sugar motifs. This leaves entire regions of chemical space, including enantiomeric mirror glycans, essentially uncharted.

Most existing predictive tools describe a glycan as a collection of monosaccharide building blocks. Models such as GLAMOUR, GNNGLY, and GIFFLAR embed glycans at the monomer level and can extrapolate to uncommon subunits, yet they struggle whenever a sugar contains a residue absent from their encoding scheme. The field of Glycomics therefore faced a persistent blind spot: how to reason about glycans whose very atoms fall outside the training vocabulary, and how to anticipate interactions of sugars that nature has rarely, if ever, synthesized.

That gap matters because recognition of glycans by glycan-binding proteins, also called lectins, sits at the heart of immunity, cell-to-cell communication, and self versus nonself discrimination. Outside the well-studied realm of mammalian N- and O-linked glycans, binding data are sparse and unevenly distributed across kingdoms. Rules for plant and bacterial glycans are assembled piecemeal through complex synthesis and specialized arrays, and archaeal N-glycans contain architectures alien to the mammalian textbook. A model that can generalize across this diversity must look past the familiar monosaccharide alphabet.

Describing Sugars Atom by Atom

Carpenter and colleagues asked a deceptively simple question: can a model predict binding using only the atomic connectivity of the glycan, ignoring the familiar monosaccharide alphabet entirely? They encoded each structure in two ways, as Morgan fingerprints and as atom q-grams, a graph-based representation that tallies local neighborhoods of atoms and their features. Intriguingly, a simple fully connected neural network performed nearly as well using these atom-level features as the authors’ prior monosaccharide-based model, GlyNet, which required explicit decomposition into polysaccharide q-grams.

When the team systematically removed individual atom features, one signal dominated all others: chirality. Representing a glycan merely as a graph whose nodes are labeled only by their handedness, R, S, or none, already recovered much of the predictive power, whereas deleting element identity or bond hybridization barely moved the result. This underscores a property rare among classical small molecules, where few chiral centers exist, glycan recognition is exquisitely sensitive to stereochemistry, which is precisely why detailed Oligosaccharide Analysis Service work is needed to capture fine structural differences that distinguish otherwise similar sugars.

The observation also revealed a curious corollary. Because glycans are so stereochemically dense, a model could reasonably predict binding even when told only the chirality pattern of each atom, with no information about which element occupied the node. This stands in sharp contrast to protein–small-molecule recognition, where the chemical identity of atoms carries most of the signal. Glycans, it seems, are read by lectins largely as a three-dimensional pattern of handedness and hydroxyl placement rather than as a list of residues.

All-atom representations of glycans feed a neural network that learns quantitative glycan–lectin binding from microarray data.

Fig. 1 All-atom representations of glycans feed a neural network that learns quantitative glycan–lectin binding from microarray data. (Carpenter, et al. 2025)

Unifying Disparate Binding Data with Fraction Bound

A practical hurdle is that binding measurements arrive in incompatible units. Glass-slide microarrays report relative fluorescent units that depend on spotting density and scanner gain, whereas solution-phase studies of glycomimetics report association constants. The authors collapsed both into a single, concentration-dependent quantity they call the fraction bound, f, the proportion of glycan occupied by a protein at a specified concentration. For the CFG microarray data, f was interpolated between the available concentrations using an array-specific minimum and a global maximum RFU; for glycomimetic affinity data, Ka values were converted via the Henderson-Hasselbalch relationship. This maneuver let them merge microarray intensities with solution affinity measurements into one coherent training set.

The resulting model, MCNet, accepts a chiral Morgan-fingerprint description of the glycan together with the protein concentration and returns a vector of fraction-bound values across lectins. Trained on the mammalian Consortium for Functional Glycomics array, MCNet reproduced known binding trends and, tellingly, flagged several glycans that the microarray had likely undercalled, suggesting the model rather than the experiment was correct about weak galectin ligands. Such computational cross-checks are increasingly valuable alongside empirical Glycan Profiling campaigns that must reconcile predicted and observed signals.

By folding concentration explicitly into the model, MCNet sidesteps a limitation of earlier approaches that predicted binding only at fixed concentrations. Binding, after all, tends toward zero as concentration tends toward zero, so any honest predictor must account for dose. Encoding concentration as an input also made it possible to resample data from different sources onto a common scale, turning what would have been incompatible experiments into a single learnable landscape.

Conversion of disparate binding measurements to a unified fraction-bound scale.

Fig. 2 Conversion of disparate binding measurements to a unified fraction-bound scale. (Carpenter, et al. 2025)

Extrapolating Beyond the Training Set

The real test was out-of-domain prediction. A model trained only on Consortium glycans (MCNet1) failed entirely on glycomimetics, and one trained only on glycomimetics (MCNet2) failed on Consortium glycans. Blending the two datasets (MCNet3), however, produced a model that handled both and improved on each alone. The lesson is that knowledge of one glycan family sharpens predictions for another, provided the training data span structurally diverse partners rather than a narrow slice of chemical space.

MCNet also reasoned about epimers, glycans differing at a single chiral center, and about enantiomers, the mirror images of known sugars. Because the combined dataset contained many epimers and diastereomers, the model learned how inverting an individual stereocenter alters recognition. This capacity points toward a future where Glycan Sequencing outputs can be fed straight into binding predictors without hand-coded rules, accelerating the exploration of sugars that are difficult or impossible to synthesize in bulk.

A further control reinforced the point. A model trained on only a subset of galectin-binding data could not extrapolate to unrelated lectins, while the broadly trained model could. The gain came not from a fancier architecture but from the breadth of stereochemical examples fed during training, a reminder that coverage of chiral diversity matters more than model depth for this class of problem.

Out-of-domain prediction performance across model variants.

Fig. 3 Out-of-domain prediction performance across model variants. (Carpenter, et al. 2025)

Mirror-Image Glycans and the Surprise of L-Glucose

The most striking prediction concerned cross-chiral recognition. By symmetry, the enantiomer of a glycan should be read very differently by a protein built from D-amino acids. MCNet correctly anticipated that mirror-image glycans would bind far worse than their natural parents, with one unexpected exception. The model forecast that L-glucose, a sugar essentially absent from nature, would bind to certain lectins that canonically recognize fucose, a prediction no training example contained.

That forecast was bold, because L-glucose is not part of any standard microarray. Yet when the team tested it experimentally using a Glycan Microarray Assay and, crucially, a Liquid Glycan Array (LiGA), binding of L-glucose to the fucose-binding lectins UEA-I and AAL was confirmed. Independent Lectin Microarray Assay measurements reproduced the result, and clonal binding assays provided additional confirmation. Structural modeling and superpositioning visualized how L-glucose nestles into the same pocket that normally accommodates fucose, with inhibition by soluble fucose further confirming shared binding sites. The agreement between a machine-learning extrapolation and multiple independent experimental and computational formats is the paper’s central validation.

The finding reframes a long-standing assumption. Cross-chiral recognition between peptides, nucleotides, and metabolites is generally poor, which is why mirror-image proteins resist degradation by natural enzymes. Glycans, however, are different: the spatial arrangement of hydroxyls can be conserved under mirror inversion for certain motifs, so a lectin tuned to fucose can be fooled by its mirror-image sugar. This plasticity makes glycan–protein recognition richer, and harder to predict, than recognition among other biopolymers.

Machine-learned MCNet anticipates cross-chiral recognition between mirror-image sugars and common lectins, validated by glycan and lectin arrays.

Fig. 4 Machine-learned MCNet anticipates cross-chiral recognition between mirror-image sugars and common lectins, validated by glycan and lectin arrays. (Carpenter, et al. 2025)

Why This Matters for Glycobiology

Cross-chiral recognition is not a laboratory curiosity. If mirror-image life forms were ever constructed, their glycans would be among the first molecules our immune system encountered, and we currently have little idea how canonical receptors would read them. Datasets that span stereochemical variation, exactly the kind assembled here, are the training ground for chirality-aware models of molecular recognition generally, a capability no other molecular family provides with such exhaustive coverage of diastereomers.

The work also reframes a foundational idea in structural biology. Anfinsen’s hypothesis holds that a protein’s folded structure is determined solely by its amino acid sequence. The authors generalize this across molecular interactions: the atomic connectivity of the two interacting partners may be necessary and sufficient to determine binding strength at a given concentration, without explicitly computing every conformation. If true, lightweight models like MCNet are not a shortcut but a principled way to read molecular recognition from connectivity alone.

For applied glycobiology, the study argues for richer, stereochemically diverse training data and for unifying measurements through a common fraction-bound scale. Platforms such as Glycoprotein Structure Analysis and complementary high-throughput arrays can supply the heterogeneous binding measurements needed to push such models further, while careful curation of enantiomeric and epimeric examples will determine how far the next generation of predictors can extrapolate.

Limitations and the Road Ahead

The authors are candid about constraints. Models trained on glass-surface arrays are largely silent for monosaccharides, which is why incorporating liquid glycan array data is flagged as a critical next step. Some predictions also failed to respect obvious boundary conditions, such as the requirement that bound fraction must fall to zero at vanishingly low concentration. And any atom absent from training, bromine for instance, becomes a cold-start problem the model cannot reason about. These gaps are less criticisms than a roadmap: future algorithms must learn the fraction bound from discrete values while respecting physical limits, and must ingest the widest possible stereochemical vocabulary.

Looking forward, the same atom-level framework could be extended to protein-aware models that take a sequence-derived protein embedding alongside the glycan fingerprint, enabling predictions for lectins never seen during training. Combined with expanding experimental arrays of mirror-image sugars, such models would finally let researchers probe the unseen mirror universe of glycans that no machine-learning system had previously considered.

Related Services & Products

Reference

  1. Carpenter, E. J., et al. (2025). Atom-level machine learning of protein-glycan interactions and cross-chiral recognition in glycobiology. Science Advances, 11, eadx6373. DOI: 10.1126/sciadv.adx6373.
Similar Posts

About Us

CD BioGlyco is a leading biotechnology company specializing in glycobiology. We deliver high-quality products and services to support cutting-edge research worldwide.

Contact Us

  • For research and manufacturing partners only. Not intended for (direct) human or veterinary use.
Copyright © CD BioGlyco. All rights reserved.