Related Experiment Video
Updated: Jun 27, 2026

Amplification of Near Full-length HIV-1 Proviruses for Next-Generation Sequencing
Published on: October 16, 2018
ViroBioTree: A Tree-Structured Biological Evidence Retrieval Framework for Viral Protein Function Annotation
Tinglian Lai1, Fuguo Liu2, Guodong Li3,4,5
1School of Mathematics and Computational Science, Guilin University of Electronic Technology, Guilin 541004, China.
None:
Accurate viral protein function annotation is essential for genomic surveillance, yet conventional retrieval-augmented generation (RAG) pipelines often fragment biological evidence into fixed-length text chunks, disrupting relationships among ORFs, annotations, structural domains, sequence motifs, residue mappings, and model-derived attention evidence. We propose ViroBioTree, a tree-structured biological evidence retrieval framework for downstream viral protein evidence review rather than a new primary annotation classifier. Built as an evidence organization layer on ViralMultiNet-derived ORF-level predictions and annotations, ViroBioTree converts sequence, annotation, structure, and attention evidence into typed biological nodes and traceable edges, then performs deterministic multi-channel recall, evidence-aware reranking, balanced TopK selection, rule-based verification, and node-cited report generation. In a demo benchmark, ViroBioTree achieved its strongest deterministic proxy performance on structure-explanation tasks, with Precision@K = 1.0, Recall@K = 1.0, and diversity = 0.52; these values reflect expected node-type and tag agreement rather than independent biological correctness. A bounded full-scale SARS-CoV-2 index contained 39,800 ORF rows, 80,000 attention records, 199,418 nodes, and 495,886 edges. In a stratified full20k diagnostic evaluation, ViroBioTree showed task-dependent advantages over LlamaIndex vector retrieval for conflict detection, evidence retrieval, and structure explanation, while LlamaIndex remained competitive or stronger for annotation-rich function annotation. A cross-family Influenza A Virus (IAV) diagnostic audit showed that the schema can represent IAV evidence namespaces while explicitly exposing missing formal ORF inputs, missing attention evidence, and unavailable residue/PDB assertions. Supplementary robustness, external sanity-check, diversity-risk, expert-evaluation, domain-tool positioning, and cross-family audit analyses supported traceability, report quality, and conservative evidence handling, but also showed that stable Precision@K under query perturbation does not necessarily imply stable retrieved evidence sets. ViroBioTree operates offline and deterministically, but does not address raw-read assembly, base calling, primary ORF prediction, or wet-lab validation. Its results should be interpreted as proxy and expert-reviewed evidence for traceable viral protein evidence retrieval and report generation rather than as direct validation of biological function annotation.
More Related Videos
Related Concept Videos
Size and Structure of Viral Genomes
Applications of Molecular Taxonomy
Viruses with RNA Genomes
Bacteriophages of the Human Virome
Phylogenetic Trees
Phylogenetic Trees

