Downloads
Machine-readable exports of the record set, the original expression-summary spreadsheets, and the vector sequence files. The 1870 per-construct files — annotated GenBank sequences, vector maps and sequence PDFs — are linked from each gene record rather than listed here.
Record exports
Generated from the same dataset that builds the pages, so they cannot drift out of step with the site.
Expression summary spreadsheets
These are the original workbooks from the legacy site, unmodified. They carry the measured expression levels for the constructs, which the individual gene records do not.
Destination vector sequences
| Vector | System | Files |
|---|---|---|
| pGEn1-DEST | Mammalian | |
| pGEn2-DEST | Mammalian | |
| pGEn3-DEST | Mammalian | |
| pGEc1-DEST | Mammalian | |
| pGEc2-DEST | Mammalian | |
| Ac-polh-CtermHisStrep-DEST | Baculovirus | |
| Ac-polh-NtermMelHisStrep-DEST | Baculovirus | |
| Ac-polh-NtermMelHisAviTagGFP-DEST | Baculovirus | |
| pET16-DEST | Bacterial | |
| pET32-DEST | Bacterial |
Protocols
Column definitions for the CSV exports
| Column | Meaning |
|---|---|
| gene_symbol | Gene symbol as recorded, including any parenthesised legacy synonym. |
| symbol | Current gene symbol with the synonym stripped. |
| synonyms | Legacy synonyms, semicolon-separated. |
| enzyme_class | GT, GH, Other or Bacterial. |
| cazy_family | CAZy family label(s) parsed from the record; blank when the enzyme is not CAZy-classified. |
| uniprot | UniProt accession. |
| gene_id | NCBI GeneID. |
| dna_refseq | RefSeq nucleotide accession. |
| protein_refseq | RefSeq protein accession. |
| mgc_acc | Mammalian Gene Collection accession of the PCR template. |
| protein_name | Protein name or GenBank annotation as recorded. |
| species | Source species; populated for the bacterial homologues. |
| domain_structure | Membrane topology class of the native enzyme. |
| truncation_eukaryotic | Truncation applied for eukaryotic hosts. |
| truncation_bacterial | Truncation applied for the bacterial host. |
| n_constructs | Number of construct entries on the record. |
| n_sequence_files | Number of linked GenBank/map/sequence files. |
| dnasu_clone_ids | All DNASU clone IDs on the record, semicolon-separated. |
| availability | dnasu, jarvis or pending — see the availability key on any family page. |
| record_page | Relative path to the record page in this site. |
Known gaps in the file set
36 construct file references on the original site pointed at files that server did not hold — requesting them returned 404, so there is nothing to migrate. Those references are kept visible on the affected records, marked unavailable, because the reference is itself evidence the construct was built; they are excluded from the CSV and JSON exports, which list only files that exist. Every other referenced file is present in this site.
all_constructs.csv) has one row per construct rather than per gene, and adds construct_kind (entry clone or expression construct), host, vector, fusion_strategy, dnasu_clone_id and files — the site-relative paths of that construct's sequence files, semicolon-separated.