Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Subset clustering of binary sequences, with an application to genomic abnormality data.

Peter D Hoff1

  • 1Department of Statistics and Biostatistics, University of Washington, Seattle, 98195-4322, USA. hoff@stat.washington.edu

Biometrics
|January 13, 2006
PubMed
Summary

This study introduces a flexible model for clustering complex binary data, identifying distinct groups and their unique characteristics. This approach is particularly useful for analyzing genomic data and understanding disease patterns.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Smaller p-values in genomics studies using distilled auxiliary information.

Biostatistics (Oxford, England)·2021
Same author

JOINT MEAN AND COVARIANCE MODELING OF MULTIPLE HEALTH OUTCOME MEASURES.

The annals of applied statistics·2019
Same author

Fast Inference for the Latent Space Network Model Using a Case-Control Approximate Likelihood.

Journal of computational and graphical statistics : a joint publication of American Statistical Association, Institute of Mathematical Statistics, Interface Foundation of North America·2016
Same author

MULTILINEAR TENSOR REGRESSION FOR LONGITUDINAL RELATIONAL DATA.

The annals of applied statistics·2016
Same author

Testing for nodal dependence in relational data matrices.

Journal of the American Statistical Association·2016
Same author

Testing and Modeling Dependencies Between a Network and Nodal Attributes.

Journal of the American Statistical Association·2016

Area of Science:

  • Statistics
  • Bioinformatics
  • Computational Biology

Background:

  • Clustering multivariate binary data presents challenges due to complex dependencies.
  • Existing methods may struggle to identify cluster-specific distinguishing attributes.

Purpose of the Study:

  • To develop a model-based clustering approach for multivariate binary data.
  • To allow cluster-specific attributes that may vary between clusters.
  • To enable unified estimation of cluster number, memberships, and parameters.

Main Methods:

  • Utilizes a multivariate Dirichlet process mixture model.
  • Employs a unified estimation framework for model parameters.
  • Applies the method to analyze genomic abnormality data.

Related Experiment Videos

Main Results:

  • The proposed model effectively clusters multivariate binary data.
  • It allows for the identification of unique attributes defining each cluster.
  • Demonstrates applicability in analyzing genomic data, such as tumor development.

Conclusions:

  • The multivariate Dirichlet process mixture model offers a powerful tool for clustering binary data.
  • This approach facilitates a deeper understanding of complex biological data, including genomic abnormalities.
  • Provides a nonparametric estimation scheme for dependent binary sequences.