检测功能排名的有影响力的点
1Medical Faculty Heidelberg, Heidelberg University, Heidelberg 69120, Germany; Institute of Medical Biometry and Statistics, Faculty of Medicine and Medical Center University of Freiburg, Freiburg 79104, Germany.
Computational biology and chemistry
|January 10, 2025
概括
影响点 (IP) 可以扭曲生物信息学特征排名. 本研究引入了一种用于检测这些知识产权的新方法,提高了下游分析的可靠性,并强调了常规知识产权检测的必要性.
科学领域:
- 生物信息学是一种生物信息学.
- 数据分析 数据分析
- 统计建模 统计建模
背景情况:
- 在生物信息学中,特征排名对于数据解释至关重要.
- 影响力点 (IP) 可以显著扭曲这些排名,往往没有被发现.
- 被忽视的知识产权可以导致不准确的生物学见解.
研究的目的:
- 调查影响力点对生物信息学特征排名的影响.
- 开发和评估一种用于检测影响力点的新方法.
- 为了强调确定IPs对于可靠的下游分析的重要性.
主要方法:
- 采用一个"留下一个"的方法来评估单个数据点的影响.
- 开发了一种新的等级比较方法,使用适应性优先权重量.
- 拟议的IP检测方法在多个公共数据集上得到了验证,包括TCGA基因表达数据.
主要成果:
- 开发的方法成功地确定了基因表达数据集中的影响点.
- 结果表明,知识产权可以大幅扭曲特征排名.
- 检测到的排名扭曲可能会对后续分析产生负面影响,例如路径丰富.
结论:
- 影响力点对功能排名和随后的生物信息学分析产生重大影响.
- 常规检测影响点至关重要,但目前未得到充分利用.
- 开发的IP检测方法作为一个名为"findIPs"的R包提供.
相关概念视频
Outliers and Influential Points
4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K
Ranks
225
Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
225
Review and Preview
6.9K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
Percentiles are a type of fractile that partition data into...
6.9K
Weighted Mean
4.9K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
4.9K
Quartile
4.1K
Quartiles are numbers that separate the data into quarters. Quartiles may or may not be part of the data. To find the quartiles, first, find the median or second quartile. The first quartile, Q1, is the middle value of the lower half of the data, and the third quartile, Q3, is the middle value, or median, of the upper half of the data. To get the idea, consider the same data set:
1; 1; 2; 2; 4; 6; 6.8; 7.2; 8; 8.3; 9; 10; 10; 11.5
The median or second quartile is seven. The lower half of the...
1; 1; 2; 2; 4; 6; 6.8; 7.2; 8; 8.3; 9; 10; 10; 11.5
The median or second quartile is seven. The lower half of the...
4.1K
Percentile
6.5K
A percentile indicates the relative standing of a data value when data are sorted into numerical order from smallest to largest. It represents the percentages of data values that are less than or equal to the pth percentile. For example, 15% of data values are less than or equal to the 15th percentile.
6.5K


