Related Experiment Video
Updated: Feb 13, 2026

08:32
Orthotopic Aortic Transplantation: A Rat Model to Study the Development of Chronic Vasculopathy
Published on: December 4, 2010
12.7K
Evaluation of Large Language Models for Peer Review in Transplantation Research: Algorithm Validation Study.
Selena Ming Shen1, Zifu Wang2, Krittika Paul3
1Pine View School, Osprey, FL, United States.
JMIR AI
|February 11, 2026
Summary
Large language models (LLMs) show promise for assisting in scientific peer review, particularly in organ transplantation. However, current LLMs lack the accuracy for independent review and require human oversight.
Area of Science:
- Medical research
- Artificial intelligence in science
Background:
- Peer review is crucial for research quality but faces challenges from reviewer fatigue and bias.
- The increasing volume of scientific publications necessitates exploring AI solutions like large language models (LLMs) for peer review support.
Purpose of the Study:
- To compare the performance of five open-source LLMs in reviewing organ transplantation papers.
- To evaluate the impact of author affiliations on LLM review outcomes.
- To examine the effectiveness of different prompt engineering strategies (zero-shot, few-shot, ToT, RAG) on LLM review decisions.
Main Methods:
- Evaluated 200 transplantation papers using five LLMs (Llama 3.3, Mistral 7B, Gemma 2, DeepSeek r1-distill Qwen, Qwen 2.5).
- Tested four prompt engineering strategies (zero-shot, few-shot, ToT, RAG) across various temperature settings.
- Assessed LLM performance on quartile categorization, considering author affiliations (prestigious, less prestigious, none) to detect bias.
Main Results:
- Retrieval-Augmented Generation (RAG) with a temperature of 0.5 yielded the best performance (0.35 exact match accuracy).
- LLMs tended to assign papers to middle quartiles (2 and 3), avoiding extremes.
- No significant affiliation bias was detected across models, though some showed marginal bias (Gemma 2, Qwen 2.5).
- Mistral demonstrated the highest accuracy (0.35) with the lowest computational cost.
Conclusions:
- Current open-source LLMs are not sufficiently accurate to replace human peer reviewers.
- While LLMs show potential for fairness by reducing affiliation bias, their accuracy limitations necessitate human supervision.
- Mistral offered the best balance of accuracy and efficiency; RAG is a promising prompting strategy.
Related Concept Videos
Language
921
Language is a unique communication system that uses words and systematic rules to organize and transmit information. Unlike other forms of communication, which may involve postures, movements, odors, or vocalizations, language relies on symbols and grammar. This makes human communication distinct from that of other species, who also communicate but do not use language in the same way humans do.
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
921
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
322
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
322
Review and Preview
8.4K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
Percentiles are a type of fractile that partition data into...
8.4K
Review and Preview
11.6K
Data are individual items of information obtained from a population or sample. Data may be classified as qualitative (categorical), quantitative continuous, or quantitative discrete. Because it is not practical to measure the entire population in a study, researchers use samples to represent the population. A random sample is a representative group from the population chosen by using a method that gives each individual in the population an equal chance of being included in the sample. Random...
11.6K
Reliability and Validity
14.1K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
14.1K
Influence of Parents and Peers on Identity
589
Adolescence is a pivotal period of identity formation, during which individuals begin to answer questions central to their sense of self, such as "Who am I?" and "Who do I hope to become?" Both parents and peers play critical roles in guiding adolescents through this complex developmental phase.
Parental Influence on Identity Development
Parents serve as primary guides and managers in an adolescent's life, offering support instrumental in decision-making and personal growth....
Parental Influence on Identity Development
Parents serve as primary guides and managers in an adolescent's life, offering support instrumental in decision-making and personal growth....
589

