KIMI: Knockoff Inference for Motif Identification from molecular sequences with controlled false discovery rate
Xin Bai1, Jie Ren1, Yingying Fan2
1Quantitative and Computational Biology Program, Department of Biological Sciences, Los Angeles, CA 90089, USA.
Bioinformatics (Oxford, England)
|October 29, 2020
Summary
KIMI, a new framework, identifies significant DNA patterns (k-mers) in microbial communities for accurate sequence classification. It controls false discovery rates, improving prediction accuracy for microbial groups like viruses and bacteria.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Metagenomic sequencing generates vast amounts of data, requiring efficient methods for classifying microbial sequences.
- Current k-mer frequency methods often use all k-mers, potentially including irrelevant ones, limiting prediction accuracy.
- Identifying biologically significant nucleotide patterns (k-mers) is crucial for distinguishing microbial groups.
Purpose of the Study:
- To develop a general framework, KIMI, for selecting relevant k-mers with guaranteed false discovery rate (FDR) control.
- To enable motif discovery for distinguishing microbial sequence groups at a user-defined FDR level.
- To improve the accuracy of metagenomic sequence classification by utilizing selected, biologically significant k-mers.
Main Methods:
- Developed KIMI, a framework utilizing the model-X Knockoffs statistical method for FDR control.
- Applied KIMI to select significant k-mers for sequence motif discovery.
- Evaluated KIMI's performance through simulation studies and real-world viral and bacterial contig datasets.
Main Results:
- KIMI effectively controls FDR while maintaining high statistical power in simulations.
- KIMI outperforms established methods like Benjamini-Hochberg and q-value in FDR control.
- Using KIMI-selected k-mers for prediction models significantly increases the accuracy of classifying viral and bacterial contigs.
Conclusions:
- KIMI provides a theoretically guaranteed and effective approach for selecting biologically relevant k-mers.
- The framework enhances the accuracy of microbial sequence classification by focusing on significant nucleotide patterns.
- KIMI offers a reproducible and powerful tool for metagenomic data analysis and motif discovery.


