确保外包数据生成的可重复性
Daniel B Sloan1, Mark D Stenglein2
1Department of Biology, Colorado State University, Fort Collins, Colorado, United States of America.
PLoS biology
|January 15, 2025
概括
研究人员必须报告在外包设施中生成的大数据的详细方法和样本元数据. 这提高了大数据研究的科学可重复性.
科学领域:
- 数据科学数据科学数据科学
- 生物技术是生物技术.
- 科学研究 科学研究
背景情况:
- 大数据研究正在迅速扩大,经常利用外包或集中设施.
- 大数据研究的一个重大挑战是缺乏全面的方法信息.
- 这种缺陷阻碍了复制和验证发现的能力.
研究的目的:
- 强调对大数据中的方法和样本元数据进行详细报告的关键需求.
- 为研究人员,服务提供商和其他利益相关者如何改善数据报告提供指导.
- 倡导标准,以提高大数据研究的科学可重复性.
主要方法:
- 本研究概述了记录实验和分析程序的最佳实践.
- 它详细介绍了样本元数据的重复性所需的基本组成部分.
- 作者提出了一个透明的数据生成和管理框架.
主要成果:
- 不完整的方法报告是大数据科学可重复性的主要障碍.
- 对方法和元数据的标准化报告可以显著提高数据的可用性和验证性.
- 采用这些实践有助于促进合作,加速科学发现.
结论:
- 实施强大的方法论和样本元数据报告标准对于大数据研究至关重要.
- 提高数据生成的透明度对于确保科学发现的可靠性和可重复性至关重要.
- 所有参与大数据生成的各方都应该优先考虑详细和准确的文档.
更多相关视频
相关概念视频
Data Validation
141
Method validation is a crucial process in analytical chemistry designed to confirm that a given method consistently produces reliable and high-quality results. This process is essential when a method is applied to different sample matrices or when procedural modifications are made, ensuring that the results meet acceptable standards across various applications.
Key parameters for method validation include:
Key parameters for method validation include:
141
Random Error
810
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
810
Systematic Error: Methodological and Sampling Errors
1.4K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
1.4K
Detection of Gross Error: The Q Test
5.6K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
5.6K
Bootstrapping
583
The term "bootstrap" originated in the 19th century as a metaphor for self-improvement or achieving something independently, without external assistance. This concept extends to statistical bootstrapping, a self-contained method for estimating population parameters through resampling, even though it can be computationally intensive. Developed by the American statistician Dr. Bradley Efron in 1979, bootstrapping provides a robust way to perform inference when the original sample size is...
583
Quantifying and Rejecting Outliers: The Grubbs Test
1.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K


