Related Experiment Video
Updated: Oct 22, 2025

12:44
Identification of Key Factors Regulating Self-renewal and Differentiation in EML Hematopoietic Precursor Cells by RNA-sequencing Analysis
Published on: November 11, 2014
12.5K
Identifying novel transcript biomarkers for hepatocellular carcinoma (HCC) using RNA-Seq datasets and machine
Rajinder Gupta1, Jos Kleinjans1, Florian Caiment2
1Department of Toxicogenomics, School of Oncology and Developmental Biology (GROW), Maastricht University, Maastricht, The Netherlands.
BMC Cancer
|August 27, 2021
Summary
Hepatocellular carcinoma (HCC) early detection is improved using RNA-Seq data and machine learning. Three novel transcript biomarkers (PARP2-202, SPON2-203, CYREN-211) show high accuracy in differentiating HCC from healthy cells.
Area of Science:
- Biochemistry
- Genomics
- Computational Biology
Background:
- Hepatocellular carcinoma (HCC) poses a significant global health challenge due to limited prognostic accuracy.
- Current diagnostic methods, including imaging and serum biomarkers, are insufficient for early detection.
- Transcriptomics offers a promising avenue for identifying more accurate early biomarkers.
Purpose of the Study:
- To leverage RNA-Seq data and machine learning (ML) for identifying novel, high-accuracy transcript biomarkers for HCC.
- To evaluate the efficacy of ML algorithms in pinpointing key molecular signatures indicative of HCC.
Main Methods:
- RNA-Seq data from healthy and HCC cell models were analyzed using five ML algorithms (Random Forest, KNN, Naïve Bayes, SVM, Neural Networks).
- Performance metrics including sensitivity, specificity, and AUC-ROC were evaluated.
- Recursive feature elimination was employed to identify the minimal set of transcripts with maximum discriminatory power.
Main Results:
- Transcriptomics data demonstrated superior efficiency over traditional protein biomarkers for HCC prognosis.
- Random Forest and Support Vector Machine algorithms exhibited the best performance.
- Three transcripts (PARP2-202, SPON2-203, CYREN-211) were identified with 0.97 accuracy and 0.93 kappa, significantly outperforming random selections.
Conclusions:
- RNA-Seq combined with ML effectively identifies novel transcript biomarkers for HCC.
- The identified biomarkers PARP2-202, SPON2-203, and CYREN-211 show high potential for early HCC detection.
- The developed ML pipeline is adaptable for discovering biomarkers in other RNA-Seq datasets.
Related Concept Videos
RNA-seq
10.6K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.6K
lncRNA - Long Non-coding RNAs
9.1K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
9.1K

