如何评估您的数据的公平性 - 测试两个公平验证器的总结
Caroline Stellmach1, Michael Rusongoza Muzoora1
1Berlin Institute of Health at Charité - Universitätsmedizin Berlin.
Studies in health technology and informatics
|January 25, 2024
概括
测试了两个网络工具,以评估基因组数据的公平性. 虽然这两种工具都提供了相似的分数,但都没有完全评估FHIR® JSON元数据,这凸显了需要改进FAIR数据验证方法的需要.
科学领域:
- 生物信息学是一种生物信息学.
- 数据科学数据科学数据科学
- 基因组学就是基因组学.
背景情况:
- 医疗保健和基因组学研究依赖于可查找,可访问,可互操作和可重复使用 (FAIR) 数据.
- 确保数据的公平性对于可靠的研究至关重要,但可能具有挑战性.
- 公平的数据验证器为评估和改善数据质量提供了潜在的解决方案.
研究的目的:
- 评估两个开放访问的网络工具F-UJI和FAIR-Checker在确定基因组数据的FAIR级别方面的有效性.
- 评估这些工具是否适用于不同的数据格式,包括JSON,TXT和CSV.
- 确定当前FAIR验证工具的局限性,特别是对于像FHIR® JSON元数据这样的专业格式.
主要方法:
- 使用了F-UJI和FAIR-Checker的演示版本.
- 三个JSON,TXT和CSV格式的基因组数据文件被提交进行分析.
- 我们比较了每个工具分配的FAIR分数和评级.
主要成果:
- 无论是F-UJI还是FAIR-Checker,对测试的基因组数据文件都产生了可比的FAIR分数.
- F-UJI提供了总的 FAIR 评分,而 FAIR 检查提供了根据 FAIR 原则分解的分数.
- 两种工具都没有熟练地评估FHIR® JSON元数据文件的公平性.
结论:
- 尽管目前的发展阶段,FAIR验证器工具在帮助研究人员进行数据FAIR化方面表现有前途.
- 需要进一步开发,以提高FAIR验证工具的功能,特别是对于复杂的数据结构,如FHIR® JSON.
- 这些工具可以帮助临床医生和研究人员提高基因组数据的质量和可重复使用性.
更多相关视频
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
771
05:51Assessing the Accuracy of Fitness Smartwatch Data for Cardiovascular and Physical Activity Monitoring: A Validation Study in Digital Health
Published on: February 21, 2025
461
相关概念视频
Data Validation
5.0K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
5.0K
Bias
4.2K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
4.2K
Friedman Two-way Analysis of Variance by Ranks
197
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
197
One-Way ANOVA: Equal Sample Sizes
3.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.3K
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
One-Way ANOVA: Unequal Sample Sizes
5.8K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.8K
