Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Bioequivalence Data: Statistical Interpretation01:16

Bioequivalence Data: Statistical Interpretation

179
Body:The statistical interpretation of bioequivalence data is a significant aspect of pharmaceutical research. Bioequivalence refers to the absence of any significant difference in the rate and extent to which the active ingredient in pharmaceutical products becomes available at the site of drug action when administered at the same molar dose under similar conditions. This helps determine if different drug products have similar absorption rates, ensuring their interchangeability.Statistical...
179
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

3.4K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.4K
Comparing Experimental Results: Student's t-Test01:09

Comparing Experimental Results: Student's t-Test

4.7K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
4.7K
Improving Translational Accuracy02:07

Improving Translational Accuracy

14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.5K
3.5K
Statistical Analysis: Overview01:11

Statistical Analysis: Overview

14.0K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
14.0K

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Electrical Impedance Tomography for Real-Time PEEP Monitoring and Atelectasis During Mask Ventilation: A Randomized Controlled Physiological Trial.

Anesthesia and analgesia·2026
Same author

EFLM position statement on the proposed 2025/0404(COD) IVDR Amendment of Article 5.5.

Clinical chemistry and laboratory medicine·2026
Same author

Unity among the units - a position paper by the DGKL.

Clinical chemistry and laboratory medicine·2026
Same author

CLEC3A-derived peptides exhibit broad-spectrum activity against <i>Candida auris</i> and clinically relevant pathogens.

Frontiers in cellular and infection microbiology·2026
Same author

From ordering to interpretation: a comprehensive framework for laboratory test indications.

Clinical chemistry and laboratory medicine·2026
Same author

Performance of DeepSeek-R1, ChatGPT (GPT-o3-mini), and Gemini 2.0 Flash on German Medical Multiple-Choice Questions: Comparative Evaluation.

JMIR formative research·2025

相关实验视频

Updated: Jan 7, 2026

Rup (RNA-seq Usability Assessment Pipeline) - Quality Control for Bulk RNA-seq Experiments in Eukaryotes
05:07

Rup (RNA-seq Usability Assessment Pipeline) - Quality Control for Bulk RNA-seq Experiments in Eukaryotes

Published on: November 7, 2025

309

聊天GPT和参考间隔:对GPT-3.5 Turbo,GPT-4和GPT-4o中可重复性的比较分析.

Annika Meyer1,2, Edgar Schömig3, Thomas Streichert2

  • 1Department of Anesthesiology and Operative Intensive Care, Faculty of Medicine and University Hospital, University Hospital Cologne, Cologne, Germany.

Frontiers in artificial intelligence
|December 29, 2025
PubMed
概括

像ChatGPT这样的大型语言模型在实验室医学中表现有前途,但在一致的参考间隔上扎. 新版本有所改进,但变化仍然存在,特别是在非标准化测试中.

关键词:
聊天GPT 聊天GPT 的意思聊天机器人 聊天机器人一致性的一致性大型语言模型参考时间间隔是参考时间间隔.可以重复的可重复性.

更多相关视频

Intraperitoneal Glucose Tolerance Test, Measurement of Lung Function, and Fixation of the Lung to Study the Impact of Obesity and Impaired Metabolism on Pulmonary Outcomes
08:30

Intraperitoneal Glucose Tolerance Test, Measurement of Lung Function, and Fixation of the Lung to Study the Impact of Obesity and Impaired Metabolism on Pulmonary Outcomes

Published on: March 15, 2018

14.6K
Pre-Implantation Genetic Testing for Aneuploidy on a Semiconductor Based Next-Generation Sequencing Platform
09:30

Pre-Implantation Genetic Testing for Aneuploidy on a Semiconductor Based Next-Generation Sequencing Platform

Published on: August 17, 2022

3.5K

相关实验视频

Last Updated: Jan 7, 2026

Rup (RNA-seq Usability Assessment Pipeline) - Quality Control for Bulk RNA-seq Experiments in Eukaryotes
05:07

Rup (RNA-seq Usability Assessment Pipeline) - Quality Control for Bulk RNA-seq Experiments in Eukaryotes

Published on: November 7, 2025

309
Intraperitoneal Glucose Tolerance Test, Measurement of Lung Function, and Fixation of the Lung to Study the Impact of Obesity and Impaired Metabolism on Pulmonary Outcomes
08:30

Intraperitoneal Glucose Tolerance Test, Measurement of Lung Function, and Fixation of the Lung to Study the Impact of Obesity and Impaired Metabolism on Pulmonary Outcomes

Published on: March 15, 2018

14.6K
Pre-Implantation Genetic Testing for Aneuploidy on a Semiconductor Based Next-Generation Sequencing Platform
09:30

Pre-Implantation Genetic Testing for Aneuploidy on a Semiconductor Based Next-Generation Sequencing Platform

Published on: August 17, 2022

3.5K

科学领域:

  • 医疗保健中的人工智能
  • 实验室医学和诊断 实验室医学和诊断
  • 临床病理学和信息学

背景情况:

  • 大型语言模型 (LLM) 为实验室医学中的快速临床咨询提供了潜力.
  • 关于LLM生成的参考间隔的一致性和临床可靠性存在不确定性,特别是没有临床背景.

研究的目的:

  • 评估来自三个ChatGPT版本 (GPT-3.5-Turbo,GPT-4,GPT-4o) 的参考区间输出的可重复性.
  • 通过使用参考区间变化作为应力测试来评估模型的一致性,当提示时省略区间信息.

主要方法:

  • 一项涉及72万6000个聊天机器人请求与标准化提示的横截面研究.
  • 分析了47个实验室参数中的246,842个参考间隔,以求一致性.
  • 使用变量系数 (CV) 和回归模型来评估变量的统计分析.

主要成果:

  • 参考区间的平均CV为26.50% (下限) 和15.82% (上限).
  • 在GPT-4和GPT-4o中,CVs明显低于GPT-3.5-Turbo.
  • 不一致的输出值得注意的是标准化不良的参数和不同的单位表达式.

结论:

  • 虽然最新的ChatGPT版本显示了改进的重复性,但诊断上仍然存在不可接受的变异性,特别是在非标准化分析物中.
  • 经过深思熟虑的快速设计,实验室实践的全球标准化,模型改进和监管监督至关重要.
  • 当前的人工智能聊天机器人应该仅限于专业使用,并接受培训,在没有提供参考间隔的情况下拒绝解释.