Related Experiment Video
Updated: Jan 15, 2026

Detecting Somatic Genetic Alterations in Tumor Specimens by Exon Capture and Massively Parallel Sequencing
Published on: October 18, 2013
Bridging data gaps: Methodological advances in extracting and analyzing genetic information from unstructured
Danny Styvens Cardona1, Juan Pablo Valencia-Arango1, Juan Pablo Gallo2
1Data Science Department, Bioscience Center - Ayudas Diagnósticas SURA, Medellín, Colombia; Omics Science Center, Bioscience Center - Ayudas Diagnósticas SURA, Medellín, Colombia; Personalized Medicine Group. Bioscience Center, Ayudas Diagnósticas SURA, Medellín, Colombia.
Introduction:
The global impact of cancer, driven by both acquired and hereditary mutations, underscores the necessity for extensive research efforts. Despite the increasing volume of genetic data, significant gaps remain in data science research, particularly in Latinos and admixed populations. This study utilizes advanced data science techniques to integrate genetic and clinical data, aiming to improve the understanding of hereditary cancer in Colombia and demonstrating the transformative potential of data-driven approaches in cancer research.
Methods:
This observational study analyzed healthcare databases from four regions and 11 cities in Colombia. Genetic data were extracted from PDF reports within SURA Colombia's Electronic Health Records (a Latin American health insurance provider) for individuals referred for hereditary cancer testing between October 2019 and November 2021. Variants in 30 genes, aligned with NCCN guidelines, were examined using Next-Generation Sequencing (NGS). Data extraction was automated using Python and R, followed by integration and analysis of genetic, clinical, and sociodemographic data using advanced data science tools hosted on Azure infrastructure. These tools enabled predictive modeling and cross-referencing to explore correlations between genetic variants and clinical outcomes.
Results:
The study included 1377 patients, with a predominance of women (92.81 %) and 63 % from the northwestern region of Colombia. The largest age group (40.37 %) was between 31 and 44 years, and 95.35 % had a personal cancer history, primarily breast cancer (75.86 %). Hereditary cancer testing revealed 145 positive results and 587 uncertain outcomes. Data science-driven analysis identified higher positivity rates in patients aged 31-44 and over 50, particularly in the northeast and central regions. Among positive results, 42.6 % included variants of uncertain significance, with 95.9 % of these patients having a personal cancer history.
Conclusion:
This study highlights the significant role of data science in analyzing hereditary cancer data. Advanced computational techniques can aid in genetic variant reclassification, uncover patterns in underrepresented populations, and inform personalized interventions for hereditary cancer management in Latin America.
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...

