Related Experiment Video
Updated: Jun 22, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
BioCoder: a benchmark for bioinformatics code generation with large language models
Xiangru Tang1, Bill Qian1, Rick Gao1
1Department of Computer Science, Yale University, New Haven, CT 06520, United States.
BioCoder, a new benchmark, evaluates large language models (LLMs) for bioinformatics code generation. Models with domain-specific knowledge and long-context capabilities perform best.
Area of Science:
- Bioinformatics
- Computational Biology
- Artificial Intelligence
Background:
- Large language models (LLMs) show promise in code generation but require domain specialization for complex tasks.
- Bioinformatics involves intricate algorithms, data operations, and domain knowledge, necessitating tailored LLM evaluation.
Purpose of the Study:
- Introduce BioCoder, a novel benchmark designed to assess LLM performance in generating bioinformatics-specific code.
- Evaluate the capabilities of various LLMs on complex bioinformatics coding tasks, including cross-file dependencies and class declarations.
Main Methods:
- Developed BioCoder using 1026 Python functions and 1243 Java methods from GitHub, plus 253 Rosalind Project examples.
- Employed topic modeling to ensure benchmark code representativeness of bioinformatics calculations.
- Utilized a fuzz-testing framework for rigorous LLM evaluation across multiple models.
Main Results:
- Evaluated models including GPT-4, GPT-3.5, StarCoder, and others, identifying key performance factors.
- Demonstrated that fine-tuning on domain-specific data (StarCoder) significantly improves benchmark performance (>15% Pass@K).
- Observed that models with long-context handling (>2600 tokens) and bioinformatics domain knowledge outperform general models.
Conclusions:
- LLMs need domain-specific knowledge and long-context understanding for effective bioinformatics code generation.
- BioCoder serves as a crucial tool for advancing LLM capabilities in specialized scientific domains.
- Future LLM development should prioritize integrating specialized knowledge and contextual awareness for scientific applications.
More Related Videos
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
09:34A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
Related Concept Videos
Improving Translational Accuracy
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Comparing Mitochondrial, Chloroplast, and Prokaryotic Genomes
Genomics
Biostatistics: Overview
Discrete variables are...
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...