Related Experiment Video
Updated: Sep 15, 2025

09:45
Detection of Copy Number Alterations Using Single Cell Sequencing
Published on: February 17, 2017
11.8K
LYCEUM: learning to call copy number variants on low-coverage ancient genomes
Mehmet Alper Yılmaz1, Ahmet Arda Ceylan1, Gun Kaynar2
1Department of Computer Engineering, Bilkent University, Ankara 06800, Türkiye.
Bioinformatics (Oxford, England)
|July 15, 2025
Summary
We developed LYCEUM, a machine learning tool for detecting copy number variants (CNVs) in ancient DNA (aDNA). LYCEUM accurately identifies CNVs in low-coverage aDNA, aiding studies on adaptation and disease susceptibility.
Area of Science:
- Genomics
- Bioinformatics
- Evolutionary Biology
Background:
- Copy number variants (CNVs) are key drivers of phenotypic variation and adaptation.
- CNVs contribute to various genetic disorders, making ancient DNA (aDNA) critical for understanding disease susceptibility.
- Detecting CNVs in degraded, contaminated, and low-coverage aDNA is challenging for conventional algorithms.
Purpose of the Study:
- To introduce LYCEUM, the first machine learning-based CNV caller specifically designed for ancient DNA.
- To address the limitations of existing CNV detection methods in low-quality and low-coverage aDNA data.
- To enable accurate CNV identification in ancient genomes for evolutionary and disease studies.
Main Methods:
- Developed LYCEUM, a novel machine learning model for aDNA CNV detection.
- Employed a two-step training strategy: pre-training on high-coverage data (1000 Genomes Project) and fine-tuning on limited high-confidence aDNA CNV calls.
- Adapted the model to accurately call CNVs from downsampled read-depth signals characteristic of aDNA.
Main Results:
- LYCEUM achieves accurate CNV detection in low-coverage ancient genomes.
- Segmental deletion calls by LYCEUM correlate with sample demographic history.
- CNV patterns identified by LYCEUM show evidence of natural selection.
Conclusions:
- LYCEUM overcomes significant challenges in aDNA CNV detection.
- The tool facilitates deeper insights into the role of CNVs in ancient population adaptation and evolution.
- LYCEUM provides a robust method for analyzing structural variations in ancient human and non-human genomes.
Related Concept Videos
Comparing Copy Number Variations and SNPs
17.9K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.9K
Genome Copying Errors
4.4K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.4K
Next-generation Sequencing
92.7K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
92.7K
Sanger Sequencing
757.5K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
757.5K
Gene Duplication and Divergence
6.3K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
6.3K
Single Nucleotide Polymorphisms-SNPs
15.9K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.9K

