大型语言模型自由响应的基于框架的定性分析:算法忠实性
Aliya Amirova1, Theodora Fteropoulli2, Nafiso Ahmed3
1Population Health Sciences, School of Life Course & Population Sciences, Faculty of Life Sciences & Medicine, King's College London, London, United Kingdom.
大型语言模型 (LLM) 可以模拟采访,但GPT-3.5缺乏用于in silico定性研究的算法忠实性,以便将其推广到人类群体. 未来的人工智能进步需要建立基于LLM的研究的有效性规范.
科学领域:
- 人工智能的人工智能
- 定性研究方法 定性研究方法
- 计算社会科学 计算社会科学
背景情况:
- 大规模的生成语言模型 (LLM) 为模拟人类对面试问题的反应提供了新的可能性.
- 定性方法传统上依赖于对自然语言采访的手动分析.
- 正在调查LLM产生的人工"参与者"的潜力,以产生可概括的见解.
研究的目的:
- 评估LLM产生的"参与者"是否可以使用定性分析方法进行研究,以产生可概括的见解.
- 为了评估LLM生成的响应的"算法忠实性",这个概念测量了输出反映人类信仰和态度的程度.
- 为了确定当前的LLM是否具有足够的算法忠实性,以便在质定性研究中被认为是有效的.
主要方法:
- 一个LLM (GPT-3.5) 被用来生成采访与"参与者"与人口统计学上与人类参与者匹配.
- 基于框架的定性分析被应用来比较人类和参与者采访之间的主题,结构和语调.
- 算法忠实性是通过将LLM输出与人类亚种群的信仰和态度进行比较来评估的.
主要成果:
- 基于框架的定性分析揭示了人类和参与者之间惊人相似的关键主题.
- 在人类和参与者之间的访谈结构和语气中观察到显著的差异.
- 在LLM生成的响应中发现了"超精度扭曲"的证据.
结论:
- 经过测试的LLM (GPT-3.5) 没有表现出足够的算法忠实性,使得in silico定性研究无法对人类群体进行概括.
- 人工智能的快速进步表明,算法忠实性可能会在未来得到改善.
- 有必要建立认识论规范来评估基于LLM的定性研究的有效性,确保代表各种生活经验.
更多相关视频
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
相关概念视频
Qualitative Analysis
For instance, group IV...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Improving Translational Accuracy
Mechanistic Models: Compartment Models in Individual and Population Analysis
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
