Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Hybrid deep learning models for fake news detection: case study on Arabic and English languages.

Frontiers in big data·2026
Same author

Mapping the technological evolution of generative AI: a patent network analysis.

Scientific reports·2025
Same author

Correction: Mobile Apps for COVID-19 Detection and Diagnosis for Future Pandemic Control: Multidimensional Systematic Review.

JMIR mHealth and uHealth·2024
Same author

Mobile Apps for COVID-19 Detection and Diagnosis for Future Pandemic Control: Multidimensional Systematic Review.

JMIR mHealth and uHealth·2024
Same author

Expression analysis of beta-secretase 1 (BACE1) enzyme in peripheral blood of patients with Alzheimer's disease.

Caspian journal of internal medicine·2019
Same author

Crocin treatment decreased pancreatic atrophy, LOX-1 and RAGE mRNA expression of pancreas tissue in cholesterol-fed and streptozotocin-induced diabetic rats.

Journal of complementary & integrative medicine·2019

Related Experiment Video

Updated: Jun 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

A pretrained biomedical large language model for Persian biomedical text mining.

Baqer M Merzah1, Tania Taami2, Salman Asoudeh3

  • 1Department of Computer Science, Faculty of Education, University of Kufa, Najaf, Iraq.

Scientific Reports
|June 12, 2026
PubMed
Summary

Large language models (LLMs) show promise in bioinformatics but need fine-tuning for complex medical questions. The BioPars benchmark and QA dataset were developed to evaluate these capabilities, especially for Persian medical queries.

Keywords:
Benchmarking LLMsBioinformatics evaluationBiomedical language modelsData-driven knowledge extractionLife sciences applications

More Related Videos

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

Related Experiment Videos

Last Updated: Jun 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

Area of Science:

  • Bioinformatics
  • Natural Language Processing
  • Artificial Intelligence

Background:

  • Large language models (LLMs) are increasingly utilized in life sciences for complex data analysis.
  • Their application in bioinformatics, particularly for medical question answering (QA), is an emerging area.
  • Existing LLMs show limitations in handling nuanced, real-world bioinformatics queries.

Purpose of the Study:

  • To introduce BioPars, a novel benchmark and QA dataset for evaluating LLMs in Persian medical contexts.
  • To assess the knowledge acquisition, synthesis, and evidence-based reasoning abilities of LLMs.
  • To compare the performance of prominent LLMs like ChatGPT, Llama, and Galactica on specialized medical QA tasks.

Main Methods:

  • Development of the BIOPARS-BENCH dataset from diverse scientific and medical sources.
  • Creation of the BioParsQA dataset with 5231 Persian medical questions and answers.
  • Evaluation of LLMs using BioPars, focusing on knowledge retrieval, interpretation, and evidence demonstration.
  • Comparative analysis using metrics such as ROUGE-L, BERTScore, MoverScore, and BLEURT.

Main Results:

  • LLMs demonstrate strong knowledge recall but struggle with higher-level reasoning and fine-grained inferences in medical QA.
  • The BioPars model achieved superior performance on BioParsQA, outperforming GPT-4 with a ROUGE-L score of 29.99.
  • The model also achieved high scores in BERTScore (90.87), MoverScore (60.43), and BLEURT (50.78), indicating robust performance in Persian medical QA.

Conclusions:

  • BioPars represents a significant contribution to evaluating LLMs in Persian medical QA, particularly for long-form answers.
  • Further fine-tuning of LLMs is crucial to enhance their capabilities for complex bioinformatics tasks.
  • The developed benchmark and model show potential for advancing AI applications in specialized medical domains.