GE-REX

Glyco-Enzyme Repository of EXpression constructs

Gateway® entry clones and mammalian, baculovirus and bacterial expression constructs for glycosyltransferases, glycoside hydrolases and glycan-modifying enzymes

Construct design

The central problem the Repository solves is producing glycosylation enzymes in soluble form. Each enzyme localises differently and most are membrane-anchored, so the truncation applied to one is not the truncation applied to another. This page sets out the topologies encountered and the fusion strategies built around them.

Enzyme domain structures and topologies

Mammalian glycosylation enzymes are generally found in membrane compartments of the secretory pathway. For glycosyltransferases the most common organisation — roughly 65% of them — is a catalytic domain facing the lumen of the ER or Golgi, held there by a single transmembrane segment. In some cases proteolysis in the lumen releases a soluble catalytic domain that continues through the secretory pathway and is secreted.

Domain organisation of glycosyltransferases: a lumenal catalytic domain anchored by a single transmembrane segment, with proteolytic release of a soluble domain in some cases.
Domain organisation of glycosyltransferases: a lumenal catalytic domain anchored by a single transmembrane segment, with proteolytic release of a soluble domain in some cases.

Most GTs carry an N-terminal transmembrane anchor (65%) that serves both as the membrane-insertion sequence for co-translational translocation and as the anchor itself; these proteins have no separate N-terminal signal sequence. Around 10% — particularly the ER glucuronosyltransferases — instead carry an N-terminal signal sequence with a C-terminal membrane anchor. A similar proportion (~12%) are multipass transmembrane proteins, mostly in the ER. The remaining ~13% have no discernible transmembrane segment and are cytosolic.

Membrane topologies of glycosyltransferases and the truncation applied to each class.
Membrane topologies of glycosyltransferases and the truncation applied to each class.

Glycoside hydrolases distribute differently. The most common organisation (~57%) is an N-terminal signal sequence followed by a C-terminal catalytic domain; most of these are lysosomal. A smaller group (17%) carries an N-terminal membrane anchor without a cleavable signal sequence and tends to reside in the Golgi. The remaining ~20% have neither signal sequence nor transmembrane segment and are cytosolic.

Membrane topologies of glycoside hydrolases.
Membrane topologies of glycoside hydrolases.

Topologies recorded in this collection

Counted from the domain-structure field of the 339 records that carry one. Labels are as entered by the curators; near-duplicate spellings have been merged, and a trailing question mark in the source (an uncertain call) has been dropped.

Domain structureRecords
Type II TMD186
N-term sig seq69
Cytosol23
C-term TMD23
Multipass20
N-term sig seq/KDEL seq5
Single internal TMD2
Single TMD in middle2
2 TMD1
Nterm SS1
multipass1
N-term SS + Cterm KTEL seq1
?? No Sig Seq1
?? No SS1
N-term and C-term TMD1
PM, no SS or TMD1
C-term and Nterm TMD1

Fusion protein strategy

For each GT or GH sequence a truncation was designed, where the topology allowed, to remove transmembrane segments and leave a catalytic domain that could be expressed in soluble form.

Truncation and fusion strategy: transmembrane segments are replaced by a signal sequence and tag cassette.
Truncation and fusion strategy: transmembrane segments are replaced by a signal sequence and tag cassette.

Every construct was built as a fusion to either an N-terminal or a C-terminal epitope/affinity tag cassette. Where an N-terminal signal anchor was truncated away, a signal sequence was appended to restore entry into the secretory pathway. All constructs carry a TEV protease cleavage site next to the appended tags, so the tags can be removed after purification.

The three tag strategies used in the mammalian expression vectors. A PDF version is linked below.
The three tag strategies used in the mammalian expression vectors. A PDF version is linked below.
Equivalent designs for the recombinant baculovirus constructs.
Equivalent designs for the recombinant baculovirus constructs.

Cloning workflow

Coding regions were amplified by PCR from Mammalian Gene Collection plasmid templates, or obtained by gene synthesis. Amplification appended Gateway® att recombination sites distal to the coding region and incorporated TEV protease cleavage sites into the N- or C-terminal extension. Amplimers were captured in pDONR221 entry vectors by BP recombination, then recombined by LR reaction into mammalian, baculovirus or bacterial destination vectors to give the construct library.

Workflow for construct preparation and analysis, from PCR capture through expression testing.
Workflow for construct preparation and analysis, from PCR capture through expression testing.