使用大型语言模型进行数据清理:对ChatGPT-4o的性能评估
Nevruz Ilhanli1, Esra Tokur Sonuvar1, Kemal Hakan Gulkesen1
1Biostatistics and Medical Informatics, Faculty of Medicine, Akdeniz University.
Studies in health technology and informatics
|July 1, 2025
概括
使用ChatGPT-4o的自动数据清理显示出有前途,对大多数变量实现高精度. 需要进一步的研究来解决局限性问题,特别是对于尿液葡萄糖等复杂数据.
科学领域:
- 数据科学数据科学数据科学
- 人工智能的人工智能
- 医疗信息学 医疗信息学
背景情况:
- 手动数据清理对于数据质量至关重要,但耗时且容易出现错误.
- 需要自动化数据清理方法来提高效率和准确性.
- 像ChatGPT这样的大型语言模型为自动化数据清理任务提供了潜力.
研究的目的:
- 评估ChatGPT-4o在自动化数据清理方面的性能.
- 评估ChatGPT-4o在不同数据变量的准确性和一致性.
主要方法:
- 使用ChatGPT-4o进行自动数据清理.
- 在性别,血红蛋白,路线和尿液葡萄糖变量上评估了清洁性能.
- 进行了三次试验,以评估一致性和确定变异.
主要成果:
- 聊天GPT-4o实现了高平均准确率:94.3% (性别),92.5% (血红蛋白),92.8% (路线).
- 对于尿液葡萄糖变量,观察到更低的准确性 (70.0%).
- 在各个试验中,在性别,血红蛋白和路线方面都注意到一致的准确性.
- 在试验中发现尿液葡萄糖的准确度有显著差异.
结论:
- 聊天GPT-4o显示了自动数据清理的巨大潜力,特别是结构化变量.
- 复杂或可变数据 (例如尿液葡萄糖) 的性能需要进一步调查和改进.
- 未来的研究应该专注于理解和减轻人工智能在数据清理中的局限性.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
592
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
2.6K
相关概念视频
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Quantifying and Rejecting Outliers: The Grubbs Test
2.1K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.1K
Detection of Gross Error: The Q Test
6.4K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.4K
Goodness-of-Fit Test
4.1K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
4.1K
