Curated glycosyltransferase protein records from model organisms, covering all 138 CAZy glycosyltransferase families
GT-CORE — the Glycosyltransferase Curated Online Resource for Enzymology — presents a curated, family-organised view of glycosyltransferase (GT) enzymes classified by the Carbohydrate-Active enZYmes (CAZy) sequence-based system, sampled across a hierarchical set of model organisms. Every record couples core identifiers — gene symbol, UniProt accession, NCBI Gene ID, source species — with curated enzymology: catalytic fold class, anomeric outcome, metal-ion dependence, oligomeric state, sugar-nucleotide donor, acceptor substrate specificity, solved structures, and both full-length and catalytic-domain protein sequences.
This is a public beta release. Please report problems, errors and corrections to moremen@uga.edu.
This is a representative set, not a comprehensive collection of CAZy GTs.
CAZy assigns millions of gene sequences to its GT families. GT-CORE distills that to 2874 curated entries chosen to represent every one of the 138 families, applied in three tiers:
- Primary. UniProt annotation level 4–5 in a listed model organism — 919 records.
- Secondary. Annotation level 4–5 in a species outside that list — 1839 records.
- Relaxed. Where a family had no qualifying member, the species and annotation filters were opened and a representative set chosen — 116 records below level 4, badged as provisional on the record page.
What is in a record
Identification
CAZy family assignment, gene symbol and synonyms, UniProt accession, NCBI Gene ID, source species and taxonomic domain, plus the UniProt annotation score that indicates evidence depth for that entry.
Function and localisation
Curated enzyme function with EC numbers where assigned, and subcellular compartment — the majority of eukaryotic GTs are type II membrane proteins of the secretory pathway.
Structure and mechanism
GT-A / GT-B / GT-C fold class, retaining versus inverting stereochemical outcome, divalent-cation dependence, oligomeric state, and links to every solved PDB structure.
Donor and acceptor specificity
Sugar-nucleotide donor with its PDB chemical component (CCD) code, acceptor substrate list, and acceptor CCD / GlycoCT identifiers distinguishing structurally verified codes from modelling identifiers.
Sequences
Full-length protein sequence and the delimited catalytic domain, each copyable or downloadable as FASTA directly from the record page.
External annotation
One-click links to UniProt, AlphaFold structure predictions, InterPro domain assignments, NCBI Gene, RCSB PDB and the corresponding CAZy family page.
How to navigate the site
- GT Families — the family directory. Every CAZy GT family present in the dataset is listed with its record count; published families link to a family index page.
- Family index page — a sortable, filterable table of all family members with the key annotation columns. Filter by taxonomic domain, fold type or mechanism, or type free text to match any column. The gene symbol links to the record.
- Record page — the complete curated dataset for one protein, grouped into identification, function, structure and mechanism, donor/acceptor specificity, and sequences, with external database links.
- Search Records — free-text search across all published records at once, with family and taxonomic-domain filters.
- Downloads — per-family and whole-dataset CSV, JSON and FASTA files.
Largest families by curated record count
All 138 families in the dataset are published, covering every one of the 2874 curated records.
Taxonomic composition of the dataset
| Taxonomic domain | Records |
|---|---|
| Eukaryota | 1716 |
| Bacteria | 1130 |
| Archaea | 14 |
| Viruses | 14 |
