Data dictionary
Every field carried on a record page, the spreadsheet column it comes from, and how completely it is populated across the 2874 records in the source dataset.
| Field | Source column | Record section | Records populated | Coverage | Notes |
|---|---|---|---|---|---|
| CAZy family | CaZy family | Identification and annotation | 2874 | 100% | CAZy sequence-based glycosyltransferase family assignment. |
| Gene symbol | Gene Symbol | Identification and annotation | 2806 | 98% | Approved or commonly used gene symbol; blank where none is assigned. |
| Synonyms / alternate names | Synonym | Identification and annotation | 2129 | 74% | Semicolon-delimited alternate gene and protein names. |
| UniProt accession | Uniprot | Identification and annotation | 2874 | 100% | UniProtKB accession; primary key for each record. |
| NCBI Gene ID | Gene_ID | Identification and annotation | 1548 | 54% | NCBI Gene identifier. |
| Source species | Species | Identification and annotation | 2874 | 100% | Source organism binomial. |
| Taxonomic domain | Domain | Identification and annotation | 2874 | 100% | Eukaryota, Bacteria, Archaea or Viruses. |
| UniProt annotation level | Annotation level | Identification and annotation | 2874 | 100% | UniProt annotation score, 1-5. |
| Enzyme function | Enzyme function | Function and localisation | 2874 | 100% | Curated activity description with EC number where assigned. |
| Subcellular location | Subcellular Location | Function and localisation | 1564 | 54% | Semicolon-delimited compartment assignments. |
| Fold type | Fold type | Structure and mechanism | 2874 | 100% | GT-A, GT-B or GT-C catalytic fold class. |
| Catalytic mechanism | Mechanism | Structure and mechanism | 2856 | 99% | Inverting or retaining stereochemical outcome. |
| Cation dependence | Cation | Structure and mechanism | 2851 | 99% | Divalent cation requirement. |
| Oligomeric state | Oligomer | Structure and mechanism | 608 | 21% | Quaternary structure / partner subunits. |
| PDB structures | PDB | Structure and mechanism | 232 | 8% | Semicolon-delimited PDB entry IDs, linked to RCSB. |
| Sugar nucleotide donor | Sugar nucleotide donor | Donor and acceptor specificity | 1511 | 53% | Activated sugar donor substrate(s). |
| Donor CCD code(s) | Donor CCD | Donor and acceptor specificity | 1454 | 51% | PDB Chemical Component Dictionary code(s) for the donor. |
| Acceptor substrate(s) | Acceptor | Donor and acceptor specificity | 2484 | 86% | Acceptor substrate(s); often a long semicolon-delimited list. |
| Acceptor CCD / GlycoCT | Acceptor GlycoCT | Donor and acceptor specificity | 108 | 4% | Acceptor CCD code or userCCD modelling identifier. |
| Full protein sequence | Full protein sequence | sequence | 2873 | 100% | Full-length protein sequence, one-letter amino-acid code. |
| Catalytic domain sequence | Catalytic domain sequence | sequence | 2800 | 97% | Sequence span assigned to the catalytic domain. |
| Curation note | Student | Curation | 0 | 0% | Free-text curation provenance note; populated for few records. |
Record category colour key
Row and badge colours are carried over from the fill colours of the source workbook, and encode how well characterised an entry is.
Controlled vocabularies
Fold type
GT-A and GT-B denote the two Rossmann-like nucleotide-binding fold classes that account for most soluble glycosyltransferases; GT-C denotes the integral-membrane, polyprenyl-phosphate dependent fold. Entries with no structural assignment are left unpopulated.
Mechanism
Inverting and retaining describe the stereochemical outcome at the anomeric carbon of the transferred sugar relative to the donor.
Annotation level
The UniProt annotation score, 1 (least) to 5 (most annotated). Some entries carry the
literal values None (no UniProt) or
Correct this entry, preserved verbatim from curation.
