MMM和MMMSynth:对异质表格数据进行集群,并生成合成数据
Chandrani Kumari1,2, Rahul Siddharthan1,2
1The Institute of Mathematical Sciences, Chennai, India.
PloS one
|April 17, 2024
概括
我们开发了用于集群和生成合成表格数据的新算法. 我们的马德拉斯混合模型 (MMM) 和MMMsynth方法改善了对异质数据集的数据分析和保护隐私的数据共享.
科学领域:
- 计算生物学是一种计算生物学.
- 数据科学是数据科学.
- 机器学习是机器学习.
背景情况:
- 表式数据集通常包含混合数据类型 (数值,顺序,分类).
- 在行内隐藏的集群结构可以影响结果变量,特别是在生物医学数据中.
- 患者保密法限制生物医学数据共享,增加了对合成数据生成的需求.
研究的目的:
- 在异构的表格数据集中引入用于集群和合成数据生成的新算法.
- 解决分析复杂数据结构和实现安全数据共享的挑战.
- 提高合成数据对机器学习模型培训的实用性.
主要方法:
- 开发了马德拉斯混合模型 (MMM),这是一个基于异质数据的预期最大化 (EM) 集群算法.
- 介绍了MMMsynth,一种合成数据生成算法,利用MMM的集群用于集群特定的数据分布建模.
- 基准MMMsynth与现有方法使用标准机器学习算法训练在合成和测试在现实世界的数据集.
主要成果:
- 马德拉斯混合模型 (MMM) 与标准算法相比,在合成异质数据中识别集群方面表现优越.
- MMM成功地在现实数据集中恢复了潜在的集群结构.
- MMMsynth的性能优于其他文献表格数据生成器,并且接近仅在真实数据上训练的模型的性能.
结论:
- 开发的MMM和MMMsynth算法为异质表格数据集中的集群和合成数据生成提供了有效的解决方案.
- 这些方法增强了复杂数据结构的分析,并为保护隐私的数据共享提供了可行的方法.
- 由MMMsynth生成的合成数据显示,它对于训练机器学习模型具有很高的实用性,弥合了合成和真实数据性能之间的差距.
更多相关视频
05:12ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
08:51Application of Unsupervised Multi-Omic Factor Analysis to Uncover Patterns of Variation and Molecular Processes Linked to Cardiovascular Disease
Published on: September 20, 2024
相关概念视频
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Construction of Frequency Distribution
7.6K
A frequency distribution table can be constructed using the steps given below.
First, make a table with two columns—one with the title of the data that needs to be organized, and the other column for frequency. [Draw a third column for tally marks if needed]. Then, take a look at the items given in the data set and decide if an ungrouped frequency distribution table or a grouped frequency distribution table would be more suitable. If there are large sets of different values, then it is...
First, make a table with two columns—one with the title of the data that needs to be organized, and the other column for frequency. [Draw a third column for tally marks if needed]. Then, take a look at the items given in the data set and decide if an ungrouped frequency distribution table or a grouped frequency distribution table would be more suitable. If there are large sets of different values, then it is...
7.6K
Bar Graph
16.4K
A bar graph is also called a bar chart and consists of bars that are separated from each other. It either uses horizontal or vertical bars to show comparisons among categories. The bars can be rectangles, or they can be rectangular boxes (used in three-dimensional plots). One axis of the graph represents the specific categories being compared, and the other axis shows a discrete value. In this graph, the length of the bar for each category is proportional to the number or percent of individuals...
16.4K
Contingency Table
2.5K
A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...
2.5K
Data: Types and Distribution
720
In biostatistics, data are the observations collected for analysis. There are two main types: parametric and non-parametric. Parametric data, which include continuous (e.g., weight) and discrete numerical data (e.g., number of tablets), assume a particular distribution pattern, often the normal distribution. Non-parametric data do not adhere to a specific distribution and typically comprise nominal (e.g., gender) and ordinal categorical data (e.g., pain scale ratings).
Distributions in...
Distributions in...
720
Applications of Life Tables
62
Life tables are versatile across various fields, providing a quantitative basis for analyzing mortality and survival rates. Whether used by demographers, actuaries, epidemiologists, or sociologists, life tables offer valuable insights into the dynamics of life and death, facilitating informed decisions in public health, insurance, conservation, and beyond. Their broad applicability highlights the interconnectedness of demographic data with practical outcomes in everyday life and strategic...
62
