Related Experiment Video
Updated: Feb 7, 2026

Mining Spatial Transcriptomics Datasets using DeepSpaceDB
Published on: September 5, 2025
A merged lung cancer transcriptome dataset for clinical predictive modeling
Su Bin Lim1,2, Swee Jin Tan3, Wan-Teck Lim4,5,6
1NUS Graduate School for Integrative Sciences & Engineering (NGS), National University of Singapore, #05-01, 28 Medical Drive, Singapore 117456, Singapore.
We created a user-friendly bioinformatics pipeline and a normalized dataset for non-small cell lung cancer (NSCLC) research. This tool simplifies large-scale genomic data analysis for biomarker discovery, making complex transcriptomic data accessible.
Area of Science:
- Genomics
- Bioinformatics
- Cancer Research
Background:
- The Gene Expression Omnibus (GEO) database offers vast transcriptomic data.
- Analyzing large-scale genomic data requires specialized bioinformatics skills, limiting accessibility.
- Challenges exist in data analysis, sharing, and visualization for non-experts.
Purpose of the Study:
- To develop an integrated bioinformatics pipeline for accessible transcriptomic data analysis.
- To create a normalized, preprocessed dataset for non-small cell lung cancer (NSCLC) meta-analysis.
- To facilitate biomarker discovery by simplifying data mining and statistical correction.
Main Methods:
- Integrated multiple open-source R packages into a bioinformatics workflow.
- Merged, normalized, and batch-corrected data from ten independent GEO datasets.
- Filtered genes with low variance and incorporated clinical metadata.
Main Results:
- Developed a normalized dataset comprising 1,118 patient-derived NSCLC samples and paired normal lung tissues.
- The dataset is preprocessed using robust statistical methods for differential expression analysis.
- The workflow enables large-scale meta-analysis without extensive data mining.
Conclusions:
- The integrated pipeline and normalized dataset enhance accessibility of large-scale genomic data.
- This resource serves as a powerful tool for early predictive biomarker discovery in NSCLC.
- The study overcomes technical barriers in transcriptomic data analysis for broader research application.
More Related Videos
Related Concept Videos
Predicting Molecular Geometry
Lung Capacity
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
End Point Prediction: Gran Plot
For potentiometric titration, the Gran plot is created by plotting...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...

