Related Experiment Video
Updated: Jun 16, 2025

11:08
Exploring the Effects of Spaceflight on Mouse Physiology using the Open Access NASA GeneLab Platform
Published on: January 13, 2019
12.2K
ChatGPT as Research Scientist: Probing GPT's capabilities as a Research Librarian, Research Ethicist, Data Generator,
Steven A Lehr1, Aylin Caliskan2, Suneragiri Liyanage3
1Cangrade, Inc., Watertown, MA 02472.
Summary
ChatGPT shows potential as a research tool, excelling as an ethicist and data simulator, but struggles with accurate referencing and predicting novel scientific data.
Area of Science:
- Psychological science research
- Artificial intelligence capabilities
- Large language models in research
Background:
- The integration of artificial intelligence (AI) into scientific research is rapidly evolving.
- Evaluating the capabilities of advanced AI models like ChatGPT (GPT-3.5 and GPT-4) is crucial for understanding their potential and limitations in scientific applications.
Purpose of the Study:
- To systematically assess the performance of GPT-3.5 and GPT-4 across key scientific research components.
- To evaluate AI's utility as a research librarian, ethicist, data generator, and novel data predictor within psychological science.
Main Methods:
- Four studies were conducted to probe AI capabilities: Research Librarian (reference generation), Research Ethicist (ethical protocol analysis), Data Generator (simulating known patterns), and Novel Data Predictor (predicting new empirical results).
- Psychological science served as the domain for testing, using fictional research protocols and analyzing AI-generated outputs against established benchmarks.
Main Results:
- GPT models exhibited significant hallucination rates for references (GPT-3.5: 36.0%, GPT-4: 5.4%), with GPT-4 showing improvement in acknowledging errors.
- GPT-4 demonstrated proficiency in identifying ethical violations (p-hacking) in research protocols (88.6% blatant, 72.6% subtle), while GPT-3.5 did not.
- Both models replicated known cultural biases when generating data, indicating capacity for simulating existing findings.
- Neither GPT-3.5 nor GPT-4 successfully predicted novel research results beyond their training data.
Conclusions:
- ChatGPT (GPT-3.5 and GPT-4) is a flawed but improving research librarian, prone to generating fictional references.
- GPT-4 shows promise as a research ethicist, capable of detecting ethical issues in scientific protocols.
- The models can generate data that mimics known patterns, useful for simulating results but not for discovering novel ones.
- Current AI models are limited in predicting new empirical data, highlighting areas for future development in scientific AI applications.
Related Concept Videos
Archival Research
16.0K
Some researchers gain access to large amounts of data without interacting with a single research participant. Instead, they use existing records to answer various research questions. This type of research approach is known as archival research. Archival research relies on looking at past records or data sets to look for interesting patterns or relationships. For example, a researcher might access the academic records of all individuals who enrolled in college within the past ten years and...
16.0K
Ethics in Research
22.9K
Today, scientists agree that good research is ethical in nature and is guided by a basic respect for human dignity and safety. However, this has not always been the case. Modern researchers must demonstrate that the research they perform is ethically sound.
22.9K

