Related Experiment Video
Updated: Aug 7, 2026

Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
Verification of the impact of differences between objective and subjective evaluation methods on the interpretation
Chiharu Kai1, Hideaki Tamori2,3, Yuta Hirono2,4
1Department of Intelligent Information Engineering, Research Promotion Unit, School of Medical Sciences, Fujita Health University, 1-98 Dengakugakubo, Kutsukake-Cho, Toyoake-City, Aichi, 470-1192, Japan. chiharu.kai@fujita-hu.ac.jp.
Purpose:
Research on vision-language models (VLMs) in the medical field has recently increased. However, while multifaceted evaluation is necessary to avoid the high risks associated with misdiagnosis, artificial intelligence (AI)-assisted mammogram report generation remains insufficient, with no studies on objective and subjective generation. We aimed to develop an AI system that generates mammogram reports and to verify the impact of differences between objective and subjective evaluation methods on the interpretation of this AI system.
Materials And Methods:
We used a public dataset consisting of mammograms and their reports, preparing question prompts and performing low-rank adaptation tuning on Qwen2.5(7B). We analyzed the Breast Imaging Reporting and Data System (BI-RADS) and findings agreement rate, Recall-Oriented Understudy for Gisting Evaluation (ROUGE), and Bilingual Evaluation Understudy (BLEU) for the objective evaluation. A breast clinician performed score-based evaluations of generated reports as subjective assessments. Finally, we analyzed samples of inconsistent objective and subjective results.
Results:
The BI-RADS agreement rate was 58.1%. Findings were accurately included in generated reports at 76.7% for mass and 81.4% for calcification. ROUGE-L F1 and overall BLEU were 0.672 and 0.542, respectively. Although ROUGE-L F1 or overall BLEU was above average, two samples received low scores in the generated report evaluation; these were considered over- and under-estimation.
Conclusion:
We developed an AI model that generates mammogram reports. In assessing this model, we found that multifaceted objective and subjective medical VLM evaluations are necessary for determining over- and under-estimation.

