Construct design
The central problem the Repository solves is producing glycosylation enzymes in soluble form. Each enzyme localises differently and most are membrane-anchored, so the truncation applied to one is not the truncation applied to another. This page sets out the topologies encountered and the fusion strategies built around them.
Enzyme domain structures and topologies
Mammalian glycosylation enzymes are generally found in membrane compartments of the secretory pathway. For glycosyltransferases the most common organisation — roughly 65% of them — is a catalytic domain facing the lumen of the ER or Golgi, held there by a single transmembrane segment. In some cases proteolysis in the lumen releases a soluble catalytic domain that continues through the secretory pathway and is secreted.

Most GTs carry an N-terminal transmembrane anchor (65%) that serves both as the membrane-insertion sequence for co-translational translocation and as the anchor itself; these proteins have no separate N-terminal signal sequence. Around 10% — particularly the ER glucuronosyltransferases — instead carry an N-terminal signal sequence with a C-terminal membrane anchor. A similar proportion (~12%) are multipass transmembrane proteins, mostly in the ER. The remaining ~13% have no discernible transmembrane segment and are cytosolic.

Glycoside hydrolases distribute differently. The most common organisation (~57%) is an N-terminal signal sequence followed by a C-terminal catalytic domain; most of these are lysosomal. A smaller group (17%) carries an N-terminal membrane anchor without a cleavable signal sequence and tends to reside in the Golgi. The remaining ~20% have neither signal sequence nor transmembrane segment and are cytosolic.

Topologies recorded in this collection
Counted from the domain-structure field of the 339 records that carry one. Labels are as entered by the curators; near-duplicate spellings have been merged, and a trailing question mark in the source (an uncertain call) has been dropped.
| Domain structure | Records |
|---|---|
| Type II TMD | 186 |
| N-term sig seq | 69 |
| Cytosol | 23 |
| C-term TMD | 23 |
| Multipass | 20 |
| N-term sig seq/KDEL seq | 5 |
| Single internal TMD | 2 |
| Single TMD in middle | 2 |
| 2 TMD | 1 |
| Nterm SS | 1 |
| multipass | 1 |
| N-term SS + Cterm KTEL seq | 1 |
| ?? No Sig Seq | 1 |
| ?? No SS | 1 |
| N-term and C-term TMD | 1 |
| PM, no SS or TMD | 1 |
| C-term and Nterm TMD | 1 |
Fusion protein strategy
For each GT or GH sequence a truncation was designed, where the topology allowed, to remove transmembrane segments and leave a catalytic domain that could be expressed in soluble form.

Every construct was built as a fusion to either an N-terminal or a C-terminal epitope/affinity tag cassette. Where an N-terminal signal anchor was truncated away, a signal sequence was appended to restore entry into the secretory pathway. All constructs carry a TEV protease cleavage site next to the appended tags, so the tags can be removed after purification.


Cloning workflow
Coding regions were amplified by PCR from Mammalian Gene Collection plasmid templates, or obtained by gene synthesis. Amplification appended Gateway® att recombination sites distal to the coding region and incorporated TEV protease cleavage sites into the N- or C-terminal extension. Amplimers were captured in pDONR221 entry vectors by BP recombination, then recombined by LR reaction into mammalian, baculovirus or bacterial destination vectors to give the construct library.

