Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reliability and Validity01:29

Reliability and Validity

12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Self-Report Tests of Personality01:22

Self-Report Tests of Personality

310
Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.
310
Accuracy and Errors in Hypothesis Testing01:13

Accuracy and Errors in Hypothesis Testing

176
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
176
Measures of Intelligence01:29

Measures of Intelligence

6.6K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
6.6K
Goodness-of-Fit Test01:16

Goodness-of-Fit Test

3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
Quality Assurance01:19

Quality Assurance

115
Quality assurance is the overarching term used to describe the activities employed to ensure the proper performance of a system. These activities can be classified into three categories: quality control, quality assessment, and internal corrective measures. Typically, these activities work cyclically: quality control is performed before and during the analysis, while quality assessment occurs during and after the investigation. Internal corrective measures are implemented based on the findings...
115

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Contrast use and radiation exposure during transcatheter aortic valve implantation according to valve design.

Kardiologia polska·2026
Same author

Simultaneous transcatheter edge-to-edge repair (TEER) for severe mitral and tricuspid regurgitation is feasible, safe, and associated with good clinical outcome.

PloS one·2026
Same author

Electrophysiological Correlates for the Detection of Haptic Illusions.

IEEE transactions on haptics·2025
Same author

Long-Term Follow-Up After Direct-Flow Transcatheter Aortic Valve Implantation: A Single Center Experience.

Catheterization and cardiovascular interventions : official journal of the Society for Cardiac Angiography & Interventions·2025
Same author

4D flow MRI-based grading of left ventricular diastolic dysfunction: a validation study against echocardiography.

European radiology·2025
Same author

Long-term structural valve deterioration after TAVI: insights from the EORP ESC Valve Durability TAVI Registry.

EuroIntervention : journal of EuroPCR in collaboration with the Working Group on Interventional Cardiology of the European Society of Cardiology·2025

Related Experiment Video

Updated: Jun 9, 2025

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
10:26

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities

Published on: September 11, 2021

3.9K

ChatGPT's quality: Reliability and validity of concept inventory items.

Stefan Küchemann1, Martina Rau2, Albrecht Schmidt3

  • 1Chair of Physics Education Research, Faculty of Physics, Ludwig-Maximilians-Universität München (LMU Munich), Munich, Germany.

Frontiers in Psychology
|October 23, 2024
PubMed
Summary

Large language models (LLMs) can generate physics concept items, but quality requires careful prompt engineering and expert review. Human oversight is crucial for effective educational assessments using AI-generated content.

Keywords:
ChatGPTconcept testitem creationlarge foundation modelslarge language modelsphysicsvalidation

More Related Videos

Development of a Virtual Reality Assessment of Everyday Living Skills
10:32

Development of a Virtual Reality Assessment of Everyday Living Skills

Published on: April 23, 2014

18.4K
Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
09:00

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education

Published on: August 16, 2024

709

Related Experiment Videos

Last Updated: Jun 9, 2025

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
10:26

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities

Published on: September 11, 2021

3.9K
Development of a Virtual Reality Assessment of Everyday Living Skills
10:32

Development of a Virtual Reality Assessment of Everyday Living Skills

Published on: April 23, 2014

18.4K
Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
09:00

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education

Published on: August 16, 2024

709

Area of Science:

  • Physics Education Research
  • Artificial Intelligence in Education

Background:

  • Large language models (LLMs) present opportunities and challenges in education.
  • Concerns exist regarding the quality and student overreliance on LLM-generated content.
  • This study evaluates LLM-generated conceptual items in physics education.

Purpose of the Study:

  • To assess the quality and characteristics of conceptual physics items generated by ChatGPT.
  • To compare AI-generated items with established concept inventories.
  • To understand the implications for educators using LLMs in assessment creation.

Main Methods:

  • Optimized prompts to generate 30 conceptual items in kinematics using ChatGPT.
  • Expert review and selection of the top 15 items.
  • Administered items alongside the Force Concept Inventory (FCI) to 172 university students.
  • Performed confirmatory factor analysis on student responses.

Main Results:

  • ChatGPT-generated items demonstrated medium difficulty and discrimination.
  • AI-generated items showed slightly lower average performance compared to the FCI.
  • Confirmatory factor analysis supported a three-factor model aligned with expert expectations.

Conclusions:

  • High-quality conceptual items can be generated by LLMs with significant prompt engineering and selection efforts.
  • AI-generated items approached the quality of human-created items but require careful vetting.
  • Human oversight and student feedback are essential for refining AI-generated assessments, especially for distractors.