在Hi-C接触矩阵上随机走路可以提高数据质量吗? 一个评估评估一个评估
1Department of Statistics, The Ohio State University, Columbus, Ohio, United States of America.
PloS one
|September 23, 2025
概括
随机步行改善Hi-C数据质量的方法,如随机步行与步骤 (RWS) 和随机步行与重启 (RWR),对于识别拓关联域 (TAD) 显示最小的好处. 研究人员在下游分析之前应该谨慎使用这些方法.
科学领域:
- 基因组学就是基因组学.
- 计算生物学 计算生物学
- 分子生物学分子生物学
背景情况:
- Hi-C和单细胞 Hi-C (scHi-C) 对于理解全基因组染色质组织至关重要,包括隔间,TAD和长距离相互作用.
- 数据质量,特别是稀疏的scHi-C数据,是一个问题,导致开发的方法,如随机步行与步骤 (RWS) 和随机步行与重启 (RWR) 的数据改进.
研究的目的:
- 描述基于随机走路的Hi-C数据方法.
- 在应用这些方法之前和之后,实证地研究RWS和RWR在识别拓相关域 (TAD) 的性能.
- 为随机步行方法的参数选择提供指导.
主要方法:
- 基于随机走路的方法的分析分析.
- 模拟研究来评估性能.
- 使用Hi-C和scHi-C数据集的真实数据应用.
- 在随机步行应用之前和之后评估TAD识别准确性.
主要成果:
- 随机步行方法,包括RWS和RWR,在下游TAD分析中几乎没有改善,即使具有最佳参数调.
- 缺乏选择调参数的实际指导方针 (例如,RWS的步数,RWR的重启概率) 阻碍了有效的应用.
- 该研究发现,在应用随机步行方法后,TAD识别准确度的增强很小.
结论:
- 基于随机走路的方法提高Hi-C数据质量和下游分析 (如TAD识别) 的有效性值得怀疑.
- 研究人员建议在分析之前考虑使用RWS和RWR时谨慎行事,因为观察到的益处有限.
- 可能需要进一步的研究来开发更有效的数据质量改善策略,用于Hi-C和scHi-C数据.
更多相关视频
04:13Using a Real-Time Locating System to Measure Walking Activity Associated with Wandering Behaviors Among Institutionalized Older Adults
Published on: February 8, 2019
7.2K
08:33A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
Published on: September 4, 2019
7.4K
相关概念视频
Wald-Wolfowitz Runs Test I
947
The Wald-Wolfowitz test, also known as the runs test, is a nonparametric statistical test used to assess the randomness of a sequence of two different types of elements (e.g., positive/negative values, successes/failures). It examines whether the order of the elements in a sequence is random or if there is a pattern or trend present. This nonparametric test applies to any ordered data despite the population and sample data distribution, even if a higher sample size is available.
The test works...
The test works...
947
Random Sampling Method
14.1K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
14.1K
Random and Systematic Errors
14.4K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
14.4K
Wald-Wolfowitz Runs Test II
527
The Wald-Wolfowitz runs test, commonly referred to as the runs test, is a nonparametric test used to assess the randomness of ordered data. The test evaluates the number of runs, which are consecutive sequences of similar elements within the data. If the number of runs is significantly higher or lower than expected, the data is considered non-random, indicating a detectable pattern or structure.
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
527
Random Error
8.5K
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
8.5K
Cluster Sampling Method
14.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
14.0K
