利用开源大型语言模型在医院工作人员调查中进行数据增强:混合方法研究
Carl Ehrett1, Sudeep Hegde2, Kwame Andre3
1Watt Family Innovation Center, Clemson University, Clemson, SC, United States.
JMIR medical education
|November 19, 2024
概括
开源大型语言模型 (LLM) 可以有效地增强用于文本分类的小型医疗数据集. 这种方法提高了分类器的性能,为医学教育和患者护理提供了隐私意识的解决方案.
科学领域:
- 医疗保健中的人工智能
- 自然语言处理自然语言处理.
- 医疗教育 技术 技术 医学教育
背景情况:
- 生成型大语言模型 (LLM) 显示了医学教育的潜力,但它们在医疗保健中用于增强小型数据集的应用,特别是隐私和成本限制,尚未得到充分探索.
- 现有的LLM应用程序通常依赖于第三方服务,限制其在敏感医疗保健环境中的使用.
研究的目的:
- 调查开源LLM在医疗保健中的文本分类任务中数据增强方面的有效性.
- 评估大型语言模型MetaAI (LLaMA) 和Alpaca等模型的性能,用于为医院工作人员调查生成合成数据.
主要方法:
- 采用了涉及数据增强和文本分类的两步过程.
- 四个开源生成的LLM被用来创建有关COVID-19流行病适应性的医院工作人员调查的合成数据.
- 然后使用三个不同的分类器LLM来对增强文本数据进行分类.
主要成果:
- 最好的性能是使用LLaMA 7B (温度0.7,100增强) 来增强数据和使用强大优化的BERT预训练方法 (RoBERTa) 来进行分类,平均AUC为0.87.
- 开源的LLM在有限的医疗保健数据集上显著提高了文本分类器的性能.
结论:
- 开源LLM为医疗保健环境中的数据增强提供了可行的解决方案,提高了文本分类的准确性.
- 该研究强调了在医疗应用中实施LLM时隐私和伦理考虑的重要性.
- 未来的研究应该探索LLMs在医学教育和患者护理中的进一步应用和优化.
相关概念视频
Improving Translational Accuracy
9.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.1K
Surveys
14.7K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
14.7K
Data Collection by Survey
6.4K
The systematic method of obtaining and analyzing accurate information of a population is called data collection. A survey is a standard method of data collection that involves collecting information from a target human population about their experience, opinion, or knowledge of a product, service, or process. The responses are recorded and interpreted. The most common survey examples are written questionnaires, face-to-face or telephonic conversations, focus groups, and electronic (e-mail or...
6.4K
Statistical Software for Data Analysis and Clinical Trials
498
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
498
Data Collection by Observations
11.7K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
11.7K
Statistical Methods for Analyzing Epidemiological Data
308
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
308


