Related Experiment Video
Updated: Aug 12, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Structured multi-level knowledge augmentation via small-to-large evidence-guided collaboration for knowledge-based
Meng Zhang1, Dayu Wu2, Wonjun Chung1
1College of Architecture and Design, Tongmyong University, Busan, Republic of Korea.
This study introduces an evidence augmentation framework to improve knowledge-based visual question answering (KB-VQA) by enhancing visual-semantic alignment and entity disambiguation for large language models (LLMs). The method boosts KB-VQA accuracy without altering LLM architecture or training.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Natural Language Processing
Background:
- Knowledge-based visual question answering (KB-VQA) necessitates integrating visual information with external knowledge for robust reasoning.
- Existing KB-VQA models struggle with aligning visual content to question intent and resolving ambiguous entity semantics, particularly for complex concepts.
Purpose of the Study:
- To propose an inference-time evidence augmentation framework designed to enhance the performance of frozen large language model (LLM)-based KB-VQA systems.
- To address the limitations of insufficient visual-semantic alignment and entity-level ambiguity in current KB-VQA approaches.
Main Methods:
- Developed a framework utilizing lightweight vision-language models to generate structured textual evidence for a frozen LLM.
- Integrated four modules: question-oriented image information extraction, entity enhancement via sub-questions, candidate-guided answer generation, and contextual exemplar retrieval.
- Focused on systematic construction, refinement, and organization of multimodal evidence prior to LLM inference, without modifying LLM architecture or training.
Main Results:
- Achieved 66.72% accuracy on the OK-VQA dataset and 69.51% on the A-OKVQA dataset, surpassing existing strong baselines.
- Demonstrated robustness, output-format reliability, and acceptable inference costs in supplementary analyses.
Conclusions:
- The proposed evidence augmentation framework effectively improves KB-VQA performance by enhancing multimodal evidence quality for frozen LLMs.
- The method offers a novel strategy for improving LLM-based VQA by focusing on pre-inference evidence construction and refinement.
Related Concept Videos
Observational Learning
Associative Learning
Classical conditioning, also known...
Purposive Learning
Self-Evaluation: Self-Enhancement and Self-Verification
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Inductive Reasoning