Related Experiment Video
Updated: May 4, 2026

Droplet Barcoding-Based Single Cell Transcriptomics of Adult Mammalian Tissues
Published on: January 10, 2019
Construction of a public CHO cell line transcript database using versatile bioinformatics analysis pipelines
Oliver Rupp1, Jennifer Becker2, Karina Brinkrolf3
1Center for Biotechnology, Bielefeld University, Bielefeld, Germany ; Cell Culture Technology, Bielefeld University, Bielefeld, Germany ; Bioinformatics and Systems Biology, Justus-Liebig-University, Giessen, Germany.
Abstract:
Chinese hamster ovary (CHO) cell lines represent the most commonly used mammalian expression system for the production of therapeutic proteins. In this context, detailed knowledge of the CHO cell transcriptome might help to improve biotechnological processes conducted by specific cell lines. Nevertheless, very few assembled cDNA sequences of CHO cells were publicly released until recently, which puts a severe limitation on biotechnological research. Two extended annotation systems and web-based tools, one for browsing eukaryotic genomes (GenDBE) and one for viewing eukaryotic transcriptomes (SAMS), were established as the first step towards a publicly usable CHO cell genome/transcriptome analysis platform. This is complemented by the development of a new strategy to assemble the ca. 100 million reads, sequenced from a broad range of diverse transcripts, to a high quality CHO cell transcript set. The cDNA libraries were constructed from different CHO cell lines grown under various culture conditions and sequenced using Roche/454 and Illumina sequencing technologies in addition to sequencing reads from a previous study. Two pipelines to extend and improve the CHO cell line transcripts were established. First, de novo assemblies were carried out with the Trinity and Oases assemblers, using varying k-mer sizes. The resulting contigs were screened for potential CDS using ESTScan. Redundant contigs were filtered out using cd-hit-est. The remaining CDS contigs were re-assembled with CAP3. Second, a reference-based assembly with the TopHat/Cufflinks pipeline was performed, using the recently published draft genome sequence of CHO-K1 as reference. Additionally, the de novo contigs were mapped to the reference genome using GMAP and merged with the Cufflinks assembly using the cuffmerge software. With this approach 28,874 transcripts located on 16,492 gene loci could be assembled. Combining the results of both approaches, 65,561 transcripts were identified for CHO cell lines, which could be clustered by sequence identity into 17,598 gene clusters.
Insights
Researchers developed new methods to assemble Chinese hamster ovary (CHO) cell transcripts, significantly expanding the available data for this key mammalian expression system used in biopharmaceutical production.
Area of Science:
- Biotechnology
- Genomics
- Molecular Biology
Background:
- Chinese hamster ovary (CHO) cells are the primary mammalian expression system for therapeutic protein production.
- Limited public data on CHO cell transcriptomes hinders biotechnological process optimization.
- Existing annotation systems (GenDBE, SAMS) are foundational but require expanded transcript data.
Purpose of the Study:
- To develop a high-quality CHO cell transcript set from extensive sequencing data.
- To establish robust pipelines for assembling and improving CHO cell transcripts.
- To create a publicly accessible platform for CHO cell genome/transcriptome analysis.
Main Methods:
- Constructed cDNA libraries from diverse CHO cell lines and culture conditions.
- Employed Roche/454 and Illumina sequencing technologies.
- Utilized de novo assembly (Trinity, Oases, CAP3) and reference-based assembly (TopHat/Cufflinks, GMAP, cuffmerge) pipelines.
Main Results:
- Assembled 28,874 transcripts from 16,492 gene loci using a combined approach.
- Identified a total of 65,561 transcripts for CHO cell lines.
- Clustered transcripts into 17,598 distinct gene clusters based on sequence identity.
Conclusions:
- The developed strategy significantly enhances the available CHO cell transcript data.
- This expanded dataset is crucial for advancing research and optimizing biopharmaceutical production using CHO cells.
- The established pipelines and data contribute to a more comprehensive CHO cell analysis platform.
More Related Videos
Related Concept Videos
Cell Lines
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...

