Related Experiment Video
Updated: Sep 28, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Known QAC determinants outperform genome language model embeddings in a leakage-aware public Listeria benzalkonium
Carlos Victor Montefusco-Pereira1
1Independent Researcher in Data Science and Artificial Intelligence in Industrial Pharmaceutics, Berlin, Germany.
Abstract:
Whole-genome sequencing and pretrained DNA models may help prioritize disinfectant-tolerance hypotheses, but phenotype prediction requires measured, isolate-linked labels and evaluation that keeps related lineages apart. We compared conventional sequence features, numerical DNA representations generated by DNABERT2, and known quaternary ammonium compound (QAC) determinants in a public benchmark of 197 Listeria monocytogenes isolates with benzalkonium chloride MIC data. Models were evaluated across five repeated lineage-grouped train, development, and test splits. The known-QAC determinant rule showed the highest balanced accuracy (mean ± SD, 0.961 ± 0.057) and ROC AUC (0.961 ± 0.057). A model combining k-mer, DNABERT2, and QAC scores reached balanced accuracy 0.915 ± 0.023, whereas the best DNABERT2-based classifier reached 0.570 ± 0.110; sparse k-mer models did not exceed the dummy baseline. Public non-Listeria studies lacked the row-level phenotype-genome linkage required for supervised modelling. An exploratory bioinformatic screen therefore identified only candidate regions for future testing, not validated tolerance predictions. Because the benchmark contained only 197 publicly linkable isolates, these findings are hypothesis-generating and require confirmation in larger, independent phenotype-linked cohorts. Within these limits, interpretable determinants were the primary validated signal and DNA-model scores were secondary evidence.
