Related Experiment Video
Updated: Jan 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
There are significant differences among artificial intelligence large language models when answering scientific
Francisco Javier Álvarez-Martínez1, Luis Esteban2, Lucas Frungillo3
1Institute of Research, Development and Innovation in Health Biotechnology of Elche (IDiBE), Universitas Miguel Hernández (UMH), Elche, Spain.
Claude 3.5 Sonnet and Gemini show promise in generating scientific responses, but other large language models (LLMs) need further development. AI evaluation frameworks and ethical considerations are crucial for responsible scientific research.
Area of Science:
- Artificial Intelligence
- Scientific Research
- Natural Language Processing
Background:
- Large language models (LLMs) are increasingly used in scientific research.
- Evaluating the accuracy and efficacy of LLMs for scientific tasks is essential.
Purpose of the Study:
- To comparatively evaluate the performance of five prominent free large language models in generating accurate scientific responses.
- To assess the impact of retrieval-augmented generation (RAG) and prompt engineering on LLM performance.
- To gauge expert reviewers' perceptions of AI utility, trustworthiness, and ethical considerations in scientific research.
Main Methods:
- Comparative evaluation of five free large language models (Claude 3.5 Sonnet, Gemini, ChatGPT 4o, Mistral Large 2, Llama 3.1 70B).
- Assessment by sixteen expert scientific reviewers based on depth, accuracy, relevance, and clarity.
- Application of retrieval-augmented generation (RAG) and prompt refinement techniques.
Main Results:
- Claude 3.5 Sonnet achieved the highest score, followed by Gemini.
- Significant variability in performance was observed among other models.
- Reviewer perceptions of AI utility and trustworthiness improved, but ethical concerns regarding transparency persisted.
Conclusions:
- LLMs like Claude 3.5 Sonnet show potential for scientific applications, while others may need further development or prompt engineering.
- Structured evaluation frameworks and ethical guidelines are necessary for responsible AI integration in science.
- Findings should be interpreted cautiously due to limited sample size and domain specificity.
Related Concept Videos
Language and Cognition
Natural and Artificial Concepts
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Introduction to Cognitive Psychology
This field emerged in the mid-20th century, following a period dominated by behaviorism, which...
Intelligence
Non-equilibrium in the Cell
