在基因组谱图中识别数学模式,与完整的SARS-CoV-2序列中的变异分类相关联
Ana Guerrero-Tamayo1, Borja Sanz Urquijo2, María-Dolores Moragues Tosantos3
1Faculty of Engineering, University of Deusto, 48007, Bilbao, Biscay, Spain. ana.guerrero@deusto.es.
Scientific reports
|December 5, 2025
概括
病毒基因组中的数学模式,如SARS-CoV-2,可以定义特征. 这项研究使用基因组谱图和转移学习来分类变异,揭示了识别的关键模式.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 之前的研究已经确定了病毒基因组中的数学模式.
- 这些模式可能决定病毒特征.
- 假设:固有的基因组数学模式决定了病毒特征.
研究的目的:
- 探索SARS-CoV-2 (严重急性呼吸系统综合征冠状病毒2) 变种分类中的数学模式.
- 使用基因组数据开发一种用于变种识别的方法.
- 研究特定核酸频率在变体分化中的作用.
主要方法:
- 从病毒序列生成基因组谱图.
- 使用预训练的卷积神经网络 (CNN) 进行两阶段的转移学习方法.
- 两步解释性用于识别重要的基因组区域和模式.
主要成果:
- 确定了特征特定SARS-CoV-2变异的独特数学模式.
- 突出了从S基因到3'UTR的基因组区域,因为它对变种识别至关重要.
- 核酸频率,特别是G和C,是关键标识符,在Omicron和VOC前的谱系中具有共同的模式.
结论:
- 数学模式与SARS-CoV-2变种分类有显著的关联.
- 这些模式代表了一个额外的基因组信息层,用于有效的病毒表征.
- 研究结果表明,在SARS-CoV-2血统中存在潜在的家族遗传联系或进化途径.
更多相关视频
11:02Detecting Somatic Genetic Alterations in Tumor Specimens by Exon Capture and Massively Parallel Sequencing
Published on: October 18, 2013
19.9K
07:15Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
11.3K
相关概念视频
Modern Molecular Taxonomy
548
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
548
Comparing Copy Number Variations and SNPs
18.5K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.5K
Evolutionary Relationships through Genome Comparisons
6.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.8K
Single Nucleotide Polymorphisms-SNPs
17.9K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.9K
Genome-wide Association Studies-GWAS
15.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.2K
