Related Experiment Video
Updated: Aug 14, 2026

Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
Subcategorisation of Data for AI Models in Healthcare: A Case Study in Mammography
Jessica E Goldring1,2, Elizabeth A Cooke1, Ruben van Engen3
1National Physical Laboratory, Teddington, Middlesex TW11 0LW, UK.
Abstract:
Background/Objectives: Accurate data subcategorising is vital for reliability and traceability in the training and validation of all artificial intelligence (AI) models. Methods: In this paper we show the complexity of clinical and technical features likely to affect the appearance and interpretation of mammography images and in turn affect the output of AI software used to aid clinical decisions. Results: Using mammography as a case study, the equitability covers screened population characteristics (e.g., women's age and ethnicity) and image acquisition key factors (e.g., brand of system, exposure factors, image processing). We examine some studies and available datasets of mammography images, summarising the metadata available. Conclusions: We recommend that, where possible, AI models are trained and evaluated using data that includes subcategories based on these features, ensuring increased equitability in the data and coverage of image heterogeneities; or, where not possible, that the subcategories for which the AI model is valid are clearly defined. Such practices can easily be implemented in a wide range of AI applications but are illustrated here with mammography. Clinical Relevance: Reliable AI holds invaluable potential for both clinical efficiency and accuracy in diagnosis. With appropriately categorised training data, a reduction in subjective assessment can be achieved, leading to trustworthy and rapid assessment.