Evaluating agentic AI for biological discovery in autonomous and copilot settings

Summary

Large language model (LLM) AI agents excel at exploring complex biological data but require human experts for guidance and synthesis. This study introduces a framework to evaluate AI