Related Experiment Video
Updated: Sep 13, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A six-tiered framework for evaluating AI models from repeatability to replaceability
Siqi Tian1, Alicia Wan Yu Lam1, Joseph Jao-Yiu Sung1
1Lee Kong Chian School of Medicine, Nanyang Technological University, Singapore.
Abstract:
Artificial intelligence (AI) is rapidly transforming biotechnology and medicine. But evaluating its safety, effectiveness, and generalizability is increasingly challenging, especially for complex generative models. Traditional evaluation metrics often fall short in high-stakes applications where reliability and adaptability are critical. We propose a six-tiered framework to guide AI evaluation across the dimensions of repeatability, reproducibility, robustness, rigidity, reusability, and replaceability. These tiers reflect increasing expectations, from basic consistency to deployment. Each is defined clearly, with actionable testing methodologies informed by literature. Designed for flexibility, the framework applies to both traditional and generative AI. Through case studies in diagnostics and medical large language models (LLM), we demonstrate its utility in fostering trustworthy, accountable, and effective AI for biomedicine, biotechnology, and beyond.
Related Concept Videos
Stereotype Content Model
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Distribution Reliability and Automation
Reliability and Validity
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Improving Translational Accuracy

