Related Experiment Video
Updated: Jul 12, 2026

G2-seq: A High Throughput Sequencing-based Technique for Identifying Late Replicating Regions of the Genome
Published on: March 22, 2018
Sparse time-varying log-ratios for longitudinal high-throughput sequencing data
Ruijin Lu1, Guoqi Yu2,3, Cuilin Zhang2,3
1Center for Biostatistics and Data Science, Washington University School of Medicine, St. Louis, MO, United States.
Abstract:
High-throughput, longitudinal omics data, such as metabolomics or microbiome profiles, present analytical challenges owing to their compositional nature and irregular observation times. Although existing approaches can address compositional or temporal aspects separately, very few are tailored to capture both properties simultaneously in a high-dimensional setting. We introduce LCoDaCoRe as a supervised learning method to identify sparse time-varying log-ratio features from longitudinal, compositional data. The proposed approach integrates functional data analysis and continuous relaxation to enable efficient feature selection from the log-ratio values. By expanding the log-transformed trajectories in their eigenspaces, LCoDaCoRe accommodates both dense and sparse sampling designs. In simulation studies, the proposed method demonstrated favorable performance in terms of predictive accuracy, selection sparsity, and precision compared to cross-sectional methods across varying correlation structures and outcome prevalence levels. Finally, we applied LCoDaCoRe to longitudinal lipidomics data from the NICHD Fetal Growth Studies and identified a highly interpretable log-ratio of triglycerides to sphingolipids that yielded more stable selection and better predictions for large-for-gestational-age births.

