Related Experiment Video
Updated: Dec 12, 2025

09:34
Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
34.4K
Seqminer2: an efficient tool to query and retrieve genotypes for statistical genetics analyses from biobank scale
Lina Yang1, Shuang Jiang2, Bibo Jiang1
1Department of Public Health Sciences, Penn State College of Medicine, Hershey, PA 17033, USA.
Bioinformatics (Oxford, England)
|August 7, 2020
Summary
Seqminer2 is a new R package that efficiently queries genetic variants in large biobanks. It offers faster data retrieval for millions of individuals and hundreds of millions of genetic variants.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Biobank-scale datasets require efficient tools for genetic variant analysis.
- Existing tools may not be optimized for the speed and scale of modern genomic data.
Purpose of the Study:
- To introduce seqminer2, a highly efficient R package for querying sequence variants.
- To improve the speed and efficiency of accessing genetic data from large biobanks.
Main Methods:
- Development of a novel variant-based index for VCF/BCF files.
- Reimplementation of support for BGEN and PLINK formats.
- Benchmarking against state-of-the-art tools like tabix.
Main Results:
- Seqminer2 demonstrates several magnitudes of speed improvement for query and retrieval.
- Enhanced performance for BGEN and PLINK file formats compared to alternatives.
- Facilitates rapid method development and data analysis in R.
Conclusions:
- Seqminer2 significantly accelerates the analysis of biobank-scale sequence data.
- The package's efficiency and format support aid researchers in R.
- Enables faster prototyping and development for genomic data analysis.

