Related Experiment Video
Updated: Mar 30, 2026

08:49
Improving Small RNA-seq: Less Bias and Better Detection of 2'-O-Methyl RNAs
Published on: September 16, 2019
8.2K
misFinder: identify mis-assemblies in an unbiased manner using reference and paired-end reads
Xiao Zhu1,2, Henry C M Leung3, Rongjie Wang4
1College of Computer Sciences and Information Engineering, Harbin Normal University, Harbin, Heilongjiang, China. zhuxiao.hit@gmail.com.
BMC Bioinformatics
|November 18, 2015
Summary
misFinder accurately identifies genome assembly errors using reference genomes and paired-end reads. This tool improves downstream analysis by distinguishing true structural variations from misassemblies, outperforming existing methods.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- High-throughput sequencing generates short reads, leading to genome assembly errors that impact downstream analysis.
- Existing error correction tools have limitations, including inability to distinguish real structural variations or producing false positives.
Purpose of the Study:
- To develop misFinder, a tool for accurate and unbiased identification and correction of genome assembly errors.
- To enhance genome assembly accuracy for improved downstream data analysis.
Main Methods:
- Combines reference genome information with aligned paired-end reads to the assembled sequence.
- Detects assembly errors and structural variations by comparing reference and assembled genomes.
- Analyzes aligned paired-end reads using coverage and insert distance features to distinguish error types and ensure high-confidence calls.
Main Results:
- misFinder accurately identifies assembly errors with high confidence and minimal miscalls.
- The tool effectively distinguishes correct assemblies corresponding to structural variations from mis-assembled sequences.
- Demonstrated superior performance over QUAST and REAPR in identifying true positive mis-assemblies with reduced false positives/negatives.
Conclusions:
- misFinder significantly improves genome assembly accuracy.
- The tool offers a robust solution for identifying and correcting assembly errors, outperforming existing methods.
- misFinder is available for free download, facilitating its use in genomic research.
Related Concept Videos
Genome Annotation and Assembly
21.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
21.5K
RNA-seq
12.5K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.5K
Mismatch Repair
7.0K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
7.0K

