Related Experiment Video
Updated: May 18, 2026

19:44
A Tactile Automated Passive-Finger Stimulator (TAPS)
Published on: June 3, 2009
A new estimator of the discovery probability
Stefano Favaro1, Antonio Lijoi, Igor Prünster
1Department of Economics and Statistics, University of Torino, Torino, Italy.
Biometrics
|October 3, 2012
Summary
This study introduces a Bayesian nonparametric method to estimate the discovery probability of new species in large genomic libraries. The approach quantifies the detection rate of rare genes and sample coverage of abundant genes.
Area of Science:
- Genomics
- Computational Biology
- Statistical Ecology
Background:
- Species sampling problems are critical in ecology and biology, influencing species richness evaluation and rare species estimation.
- Genomic applications face unique sampling challenges due to massive populations (genomic libraries) and limited sequencing.
- Existing methods struggle with the scale and complexity of genomic data, necessitating flexible approaches.
Purpose of the Study:
- To develop a Bayesian nonparametric approach for inferring species sampling issues in genomic contexts.
- To predict the discovery probability of new species (genes) from additional sample sizes in genomic libraries.
- To provide a method for quantifying the detection rate of rare genes and the coverage of abundant genes.
Main Methods:
- Utilized a Bayesian nonparametric framework for flexibility in handling large genomic datasets.
- Developed a novel estimator for the probability of detecting species with specific frequencies in an enlarged sample.
- Derived a closed-form expression for the estimator, allowing exact evaluation.
Main Results:
- The proposed estimator accurately predicts the discovery probability in genomic libraries.
- Quantified the rate of rare gene detection and the sample coverage of abundant genes as sample size increases.
- Demonstrated the method's efficacy using two expressed sequence tags (EST) datasets.
Conclusions:
- The Bayesian nonparametric approach offers a robust solution for species sampling problems in genomics.
- The derived estimator provides valuable insights into gene discovery and library characterization.
- This method aids in understanding rare gene detection and overall genomic library coverage.
Related Concept Videos
What are Estimates?
It isn't easy to measure a parameter such as the mean height or the mean weight of a population. So, we draw samples from the population and calculate the mean height or mean weight of the individuals in the sample. This sample data acts as a representative measure of the population parameter. These sample statistics are known as estimates.
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such as the mean,...
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such as the mean,...
Testing a Claim about Population Proportion
A complete procedure for testing a claim about a population proportion is provided here.
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Kaplan-Meier Approach
The Kaplan-Meier estimator is a non-parametric method used to estimate the survival function from time-to-event data. In medical research, it is frequently employed to measure the proportion of patients surviving for a certain period after treatment. This estimator is fundamental in analyzing time-to-event data, making it indispensable in clinical trials, epidemiological studies, and reliability engineering. By estimating survival probabilities, researchers can evaluate treatment effectiveness,...
Prediction Intervals
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
P-value
P-value is one of the most crucial concepts in statistics.
P-value stands for the probability value. P-value is the probability that, if the null hypothesis is true, the results from another randomly selected sample will be as extreme or more extreme as the results obtained from the given sample.
A large P-value calculated from the data indicates to not reject the null hypothesis. But a higher P-value does not mean that the null hypothesis is true. The smaller the P-value, the more unlikely...
P-value stands for the probability value. P-value is the probability that, if the null hypothesis is true, the results from another randomly selected sample will be as extreme or more extreme as the results obtained from the given sample.
A large P-value calculated from the data indicates to not reject the null hypothesis. But a higher P-value does not mean that the null hypothesis is true. The smaller the P-value, the more unlikely...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...

