Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional

Justin T Reese1,2, Leonardo Chimirri2,3, Yasemin Bridges2,4

  • 1Division of Environmental Genomics and Systems Biology, Lawrence Berkeley National Laboratory, Berkeley, CA, USA.