A Statistical Approach to Correcting Cross-Annotations in a Metagenomic Functional Profile Generated by Short Reads

Ruofei Du1,2, Donald Mercante1, Lingling An2

  • 1Biostatistics Program, School of Public Health, Louisiana State University Health Sciences Center, New Orleans, Louisiana, USA.

Journal of Biometrics & Biostatistics
|May 2, 2018
PubMed
Summary

This study introduces Probabilistic Latent Semantic Analysis to accurately profile metagenomic samples by correcting cross-annotation errors in short sequencing reads. This improves functional profiling and downstream analyses.

Related Concept Videos

Profile Leveling and Cross Sections01:26

Profile Leveling and Cross Sections

Profile leveling and cross-sections are surveying methods used to determine and document terrain elevations for infrastructure projects such as highways, railroads, canals, and pipelines. These methods provide data for earthwork planning and alignment of proposed routes.  Profile leveling involves measuring elevations along a fixed line to create a vertical terrain profile. A surveyor sets up a leveling instrument at the benchmark (BM) and records a backsight (BS) to determine the...
1.7K
Statistical Significance01:50

Statistical Significance

Once data is collected from both the experimental and the control groups, a statistical analysis is conducted to find out if there are meaningful differences between the two groups. A statistical analysis determines how likely any difference found is due to chance (and thus not meaningful). In psychology, group differences are considered meaningful, or significant, if the odds that these differences occurred by chance alone are 5 percent or less. Stated another way, if we repeated this...
22.2K
Genome Annotation and Assembly03:36

Genome Annotation and Assembly

The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
21.1K
Cross-Sectional Research01:50

Cross-Sectional Research

In cross-sectional research, a researcher compares multiple segments of the population at the same time. If they were interested in people's dietary habits, the researcher might directly compare different groups of people by age. Instead of following a group of people for 20 years to see how their dietary habits changed from decade to decade, the researcher would study a group of 20-year-old individuals and compare them to a group of 30-year-old individuals and a group of 40-year-old...
12.7K
Uncertainty in Measurement: Reading Instruments02:46

Uncertainty in Measurement: Reading Instruments

Counting is the type of measurement that is free from uncertainty, provided the number of objects being counted does not change during the process. Such measurements result in exact numbers. By counting the eggs in a carton, for instance, one can determine exactly how many eggs are there in the carton. Similarly, the numbers of defined quantities are also exact. For example, 1 foot is exactly 12 inches, 1 inch is exactly 2.54 centimeters, and 1 gram is exactly 0.001 kilograms. Quantities...
53.6K
Distance Corrections01:15

Distance Corrections

To achieve precise distance measurements, especially in surveying and construction, certain corrections must be applied to account for potential sources of error like the standardization errors, temperature variations, and slope adjustments.Standardization error emerges when measurement equipment undergoes changes, such as wear, repairs, or weather impacts. To address this, surveyors compare the equipment’s readings to a standard. This process identifies any deviation that might lead to...
299