Bioinformatics and artificial intelligence

The BioF:GREAT initiative integrates cutting-edge Artificial Intelligence (AI) and Machine Learning (ML) techniques to accelerate discovery in glycoscience. By leveraging modern computational models, we enhance prediction, interpretation, and visualization of glycoenzymatic functions across the Tree of Life.

Deep learning for glycoenzyme classification and function prediction

  • Mapping the glycosyltransferase fold landscape using interpretable deep learning — Nature Communications.
  • Application of protein language models to predict glycoenzyme function and substrate specificity from sequence.
  • Classification of glycosyltransferase families and folds across the Tree of Life.

Tools for data mining and visualisation

  • Protein Embedding Visualizer — a web application for exploring protein sequence space. Sequences are submitted as FASTA, optionally with a metadata table and PDB or mmCIF structures, and embedded using a protein language model; the resulting vectors are projected by UMAP, t-SNE or PCA and clustered, giving an interactive map in which related proteins group together and can be coloured by taxonomy, family or any supplied annotation. Structures, where provided, contribute per-residue structural tokens alongside the sequence. Developed in Dr. Natarajan Kannan’s laboratory at the University of Georgia (esbg.bmb.uga.edu/embedding).
  • GTXplorer — a platform for evolutionary and functional analysis of GT-A fold glycosyltransferases (uga-gta.netlify.app).
  • Visual query builder for knowledge-graph mining across glycoscience datasets.
  • JAAG — generation of AlphaFold 3 input for glycan modelling with correct stereochemistry.
Bioinformatics analysis figure
Bioinformatics analysis figure
Collaborations. Computational projects, including the Protein Embedding Visualizer above, are led from the Kannan laboratory and run in collaboration with the experimental BioF:GREAT groups — see internal projects for current work.