Related Experiment Video
Updated: May 21, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
An astronomical question answering dataset for evaluating large language models
Jie Li1, Fuyong Zhao1, Panfeng Chen1
1State Key Laboratory of Public Big Data, College of Computer Science and Technology, Guizhou University, Guiyang, 550025, China.
Researchers developed Astro-QA, a new benchmark dataset for evaluating large language models (LLMs) in astronomy question answering (QA). This dataset helps assess LLM capabilities in astrophysics and related fields.
Area of Science:
- Astronomy and Astrophysics
- Artificial Intelligence
- Natural Language Processing
Background:
- Large language models (LLMs) show promise in question answering (QA), but evaluating their astronomical knowledge is hindered by a lack of specialized benchmark datasets.
- Assessing LLM performance in astronomy requires a dataset that covers diverse subfields and question types.
Purpose of the Study:
- To introduce Astro-QA, the first comprehensive benchmark dataset for astronomical QA.
- To provide a standardized tool for evaluating LLM performance in astronomy.
- To facilitate future research and development of LLMs for astronomical applications.
Main Methods:
- Construction of Astro-QA dataset with 3,082 questions in English and Chinese, covering astrophysics, astrometry, celestial mechanics, history of astronomy, and astronomical techniques.
- Development of DGscore, a novel metric integrating objective and subjective question assessments with difficulty weighting.
- Validation of the dataset through extensive experiments with 27 open-source and commercial LLMs.
Main Results:
- The Astro-QA dataset effectively serves as a benchmark for evaluating LLMs in astronomical QA.
- Experiments demonstrate the dataset's utility in assessing LLM instruction following, knowledge reasoning, and natural language generation.
- The DGscore provides an accurate measure of LLM QA performance in the astronomical domain.
Conclusions:
- Astro-QA is a reliable benchmark for assessing LLM capabilities in astronomy.
- The dataset and DGscore can guide progress in developing specialized LLMs for astronomical research.
- This work bridges the gap in evaluating LLMs for scientific domains like astronomy.
Related Concept Videos
Improving Translational Accuracy
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Quantifying and Rejecting Outliers: The Grubbs Test
Distributions to Estimate Population Parameter

