Recovering incomplete data using Statistical Multiple Imputations (SMI): a case study in environmental chemistry
Theresa G Mercer1, Lynne E Frostick, Anthony D Walmsley
1Centre for Adaptive Science and Sustainability, Department of Geography, University of Hull, HU6 7RX, United Kingdom. T.Mercer@hull.ac.uk
Talanta
|October 4, 2011
Summary
Statistical Multiple Imputation (SMI) successfully recovered incomplete environmental chemistry data. This method enabled analysis of leached arsenic, chromium, and copper without altering data variance.
Area of Science:
- Environmental Chemistry
- Statistical Analysis
- Environmental Science
Background:
- Environmental chemistry data often contains missing values and values below the limit of detection.
- These data limitations hinder traditional statistical analysis, particularly in environmental leaching studies.
- Analyzing leached contaminants like arsenic, chromium, and copper requires robust statistical methods.
Purpose of the Study:
- To present a statistical technique for analyzing environmental chemistry data with missing values and detection limit issues.
- To demonstrate the application of Statistical Multiple Imputation (SMI) on environmental leaching data.
- To assess the impact of SMI on data variance and analytical feasibility.
Main Methods:
- Application of Statistical Multiple Imputation (SMI) to address missing data points.
- Analysis of leachate from lysimeters containing treated and untreated wood waste.
- Quantification of arsenic (As), chromium (Cr), and copper (Cu) concentrations using ICP-OES.
Main Results:
- SMI successfully recovered an incomplete dataset from an environmental leaching study.
- The re-analyzed complete dataset allowed for successful statistical assessment.
- It was demonstrated that SMI did not significantly affect the data's variance.
Conclusions:
- Statistical Multiple Imputation (SMI) is an effective method for handling missing data in environmental chemistry.
- SMI facilitates the analysis of complex environmental datasets previously hindered by data gaps.
- This technique enhances the reliability and scope of environmental contaminant studies.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least squares (OLS)...
Contaminants and Errors
Effective sample preparation is crucial for accurate and reliable laboratory analysis. During this process, two significant sources of error can arise: concentration bias from improper sample splitting and contamination caused by methods used to reduce particle size, such as grinding or homogenization. Identifying and minimizing these potential errors is crucial to ensuring the validity of the analysis.
Another key consideration is determining the appropriate number of samples required to...
Another key consideration is determining the appropriate number of samples required to...
Sampling Plans
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Censoring Survival Data
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different reasons...
Statistical Methods for Analyzing Epidemiological Data
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:


