Related Experiment Video
Updated: Jun 5, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
502
Establishing best practices in large language model research: an application to repeat prompting
Robert J Gallo1,2, Michael Baiocchi3, Thomas R Savage4
1Center for Innovation to Implementation, VA Palo Alto Health Care System, Menlo Park, CA 94025, United States.
Journal of the American Medical Informatics Association : JAMIA
|December 10, 2024
Summary
Establishing best practices in large language model (LLM) research is crucial. Ignoring correlations in repeated prompting data can lead to erroneous conclusions about model bias, as demonstrated in this study.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Scientific Research Methodology
Background:
- Large language models (LLMs) are increasingly used in research.
- Ensuring the reliability and validity of LLM research is paramount.
- Repeat prompting is a common technique in LLM studies, but its statistical implications require careful consideration.
Purpose of the Study:
- To highlight the importance of establishing best practices in large language model research.
- To illustrate how ignoring correlations in repeated prompting data can affect study outcomes.
- To demonstrate the critical need for appropriate statistical methods in LLM analysis.
Main Methods:
- Utilized data from a previous study on potential model bias in medical abstract peer review.
- Compared statistical methods that ignore output correlation from repeated prompting with a random effects method that accounts for it.
- Analyzed intraclass correlation coefficient (ICC) to quantify within-group correlation.
Main Results:
- A high within-group correlation (ICC=0.69) was observed with repeated prompting.
- Ignoring this correlation led to over a 100-fold inflation of the effective sample size.
- Accounting for correlation reversed the study's findings from significant evidence of bias to no evidence of bias.
Conclusions:
- The establishment of best practices for large language model research is urgently needed.
- Accurate analysis of LLM research requires accounting for the correlation inherent in repeated prompting.
- Failure to implement best practices can critically impact the validity of research conclusions.
Related Concept Videos
Modeling in Therapy
48
Modeling, a key technique in therapy, uses observational learning to help clients acquire and practice new skills by watching therapists demonstrate desired behaviors. This approach, rooted in Albert Bandura's concept of vicarious learning, plays a significant role in therapeutic interventions for various psychological conditions, including social anxiety, ADHD, and depression.
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
48
Language Development
317
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
317
Behaviorism
2.2K
The field of behaviorism was pioneered by figures such as Ivan Pavlov, John B. Watson, and B.F. Skinner fundamentally shifted the focus of psychology to the observable and controllable aspects of human and animal behavior. This shift marked a critical evolution in the discipline, emphasizing scientific rigor and experimental methodology.
The core premise of behaviorism is its focus on observable behavior rather than internal thoughts or feelings. This approach argues that true scientific...
The core premise of behaviorism is its focus on observable behavior rather than internal thoughts or feelings. This approach argues that true scientific...
2.2K
Elaborative Rehearsals
77
Elaborative rehearsal is a crucial cognitive strategy that strengthens information encoding in long-term memory by making meaningful connections between new data and pre-existing knowledge. This approach contrasts with maintenance rehearsal, which involves simple repetition without delving into the significance of the information. While maintenance rehearsal might temporarily keep information active in short-term memory, it is less effective for long-term retention.
The effectiveness of...
The effectiveness of...
77
Operant Conditioning Intervention
43
Operant conditioning serves as a foundational principle in therapeutic interventions aimed at modifying maladaptive behaviors. Central to this approach is the notion that behaviors, both adaptive and maladaptive, are learned through reinforcement. By analyzing the environmental factors that reinforce problematic behaviors, clinicians can design interventions to weaken these reinforcements and replace maladaptive behaviors with healthier alternatives.
In operant conditioning, behaviors that are...
In operant conditioning, behaviors that are...
43
Behavior Modification
125
Behavioral approaches have often been criticized for ignoring mental processes and focusing solely on observable behavior. However, these approaches provide an optimistic perspective for individuals seeking to change their behaviors. Rather than concentrating on intrinsic personality traits, behavioral approaches suggest that even longstanding habits can be modified by changing the reward contingencies that maintain them.
A real-world application of operant conditioning principles is applied...
A real-world application of operant conditioning principles is applied...
125

