Related Experiment Video
Updated: Jul 9, 2026

Competitive Genomic Screens of Barcoded Yeast Libraries
Published on: August 11, 2011
Penalized likelihood for sparse contingency tables with an application to full-length cDNA libraries.
Corinne Dahinden1, Giovanni Parmigiani, Mark C Emerick
1Seminar für Statistik, ETH Zürich, CH-8092 Zürich, Switzerland. dahinden@stat.math.ethz.ch
We developed new methods for analyzing complex interactions in biological data using log-linear models. Our approach efficiently handles sparse data, enabling better understanding of gene splicing and other biological networks.
Area of Science:
- Computational biology
- Systems biology
- Statistical genetics
Background:
- Joint analysis of categorical variables is crucial for systems biology.
- Log-linear models traditionally capture complex interactions but struggle with high-dimensional, sparse biological data.
- Analyzing full-length cDNA libraries of alternatively spliced genes requires robust methods for interaction detection.
Purpose of the Study:
- To develop and evaluate methods for model selection and parameter estimation in log-linear models for sparse contingency tables.
- To address challenges in analyzing high-dimensional biological data with potential zeros in contingency tables.
- To identify complex interaction patterns among categorical variables in biological networks.
Main Methods:
- Developed an efficient l1-penalization approach, extending the Lasso algorithm for log-linear models.
- Proposed regularization methods for parameter estimation and model selection in sparse contingency tables.
- Compared the proposed methods against other procedures using simulation studies.
Main Results:
- The proposed l1-penalization method efficiently handles sparse contingency tables common in computational biology.
- Demonstrated the effectiveness of the regularization methods on contingency tables from full-length cDNA libraries.
- Showcased the ability to perform model selection and parameter estimation in challenging biological datasets.
Conclusions:
- Regularization methods successfully detect complex interaction patterns among categorical variables.
- The developed algorithms are applicable to a broad range of biological problems involving categorical data analysis.
- Provides a robust framework for analyzing interactions in systems biology and genetic studies.
Related Concept Videos
Contingency Table
Introduction to Test of Independence
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
Probability Laws
Expected Frequencies in Goodness-of-Fit Tests
Determination of Expected Frequency
Friedman Two-way Analysis of Variance by Ranks

