对不同距离连接方法进行分析,以集群基因表达数据和观察性:实证研究
Joydhriti Choudhury1, Faisal Bin Ashraf1
1Brac University, Dhaka, Bangladesh.
JMIR bioinformatics and biotechnology
|June 27, 2024
概括
这项研究确定了生物数据的最佳聚类方法,找到与平均或病房链接产生高质量的基因集群的最大距离. 这有助于发现与癌症等疾病相关的类基因.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 庞大的生物数据需要确定基因与疾病的联系.
- 聚类是分析物种和基因之间的关系的关键.
- 选择最佳的距离链接指标对于各种生物数据集至关重要.
研究的目的:
- 确定最佳的距离链接策略,以可靠地聚类各种生物数据.
- 确定各种瘤背后的常见基因,以观察类效应.
主要方法:
- 评估了四种连接方法 (单个,完整,平均,病房) 和三种距离指标 (欧几里德,最大,曼哈顿).
- 使用结合轮宽度和集群内部距离的健身功能评估集群质量.
- 利用基因丰富分析来验证发现.
主要成果:
- 最大距离度量产生了最高质量的集群.
- 对于中等数据集,平均链接是最优的;对于大数据集,病房链接是最好的.
- 组合聚类并没有改善结果;确定了与三种癌症相关的类基因.
结论:
- 研究生物数据的各种聚类技术的准确性.
- 强调准确的聚类对基因疾病关联研究的重要性.
- 结果可以为未来的研究提供类似的生物数据集.
更多相关视频
相关概念视频
Epistasis Analysis
5.0K
Although Mendel chose seven unrelated traits in peas to study gene segregation, most traits involve multiple gene interactions that create a spectrum of phenotypes. When the interaction of various genes or alleles at different locations influences a phenotype, this is called epistasis. Epistasis often involves one gene masking or interfering with the expression of another (antagonistic epistasis). Epistasis often occurs when different genes are part of the same biochemical pathway. The...
5.0K
Dihybrid Crosses
74.7K
Overview
74.7K
Genome-wide Association Studies-GWAS
13.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.3K
DNA Microarrays
17.3K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
17.3K
Chromatin Position Affects Gene Expression
23.3K
Chromatin is the massive complex of DNA and proteins packaged inside the nucleus. The complexity of chromatin folding and how it is packaged inside the nucleus greatly influences access to genetic information. Generally, the nucleus' periphery is considered transcriptionally repressive, while the cell's interior is considered a transcriptionally active area.
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the...
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the...
23.3K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K


