Related Experiment Video
Updated: May 11, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Arch-Eval benchmark for assessing chinese architectural domain knowledge in large language models.
Jie Wu1, Mincheng Jiang1, Juntian Fan1
1Key Laboratory of Disaster Prevention and Structural Safety of Ministry of Education, Guangxi Key Laboratory of Disaster Prevention and Engineering Safety, College of Civil Engineering and Architecture, Guangxi University, 530004, Nanning, China.
This study introduces Arch-Eval to assess Large Language Models (LLMs) in architecture. While LLMs show varied performance, Chain-of-Thought evaluation is more accurate but slower than Answer-Only.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Construction Technology
Background:
- Large Language Models (LLMs) are increasingly used in Natural Language Processing (NLP).
- There is a lack of evaluation studies for LLMs specifically within the construction and architectural domains.
- Assessing domain-specific knowledge processing in LLMs is crucial for their practical application.
Purpose of the Study:
- To introduce the "Arch-Eval" framework for evaluating LLMs in architecture.
- To assess the performance of 14 different LLMs on architectural knowledge.
- To analyze LLM response stability and accuracy in the architectural domain.
Main Methods:
- Development of the "Arch-Eval" framework.
- Creation of a standardized dataset with at least 875 questions.
- Testing 14 LLMs over seven iterations, measuring Stability and Accuracy.
- Comparison of Chain-of-Thought (COT) and Answer-Only (AO) evaluation methods.
Main Results:
- Significant performance differences were observed among the 14 evaluated LLMs.
- The average accuracy difference between COT and AO evaluations was less than 3%.
- COT evaluation response time was substantially longer (26x) than AO (62.23s vs. 2.38s per question).
Conclusions:
- The "Arch-Eval" framework provides a reliable method for assessing LLMs in architecture.
- LLM performance in architectural knowledge question-answering varies significantly.
- Future research should focus on domain customization, reasoning, and multimodal interaction for LLMs in construction.
Related Concept Videos
Spanning Openings in Brick Walls
Lintels are primary supports used to span openings and can be crafted from materials such as reinforced concrete, steel-reinforced brick masonry, or simple steel angles. These are straightforward to install and are typically concealed...
Structural Classification of Joints
A fibrous joint is where the adjacent bones are united by fibrous connective...
Design Example: Dimensioning of Concrete Masonry Construction
The site engineer has laid out a plan for the storeroom with external dimensions of twelve feet in length and...
Archival Research
Typical Model Studies
Improving Translational Accuracy

