Outlier analyses of the Protein Data Bank archive using a probability-density-ranking approach

Chenghua Shao1,2, Zonghong Liu3, Huanwang Yang1

  • 1RCSB Protein Data Bank, Rutgers, The State University of New Jersey, Piscataway, NJ 08854, USA.

Scientific Data
|December 12, 2018
PubMed

Related Concept Videos

Archival Research01:40

Archival Research

Some researchers gain access to large amounts of data without interacting with a single research participant. Instead, they use existing records to answer various research questions. This type of research approach is known as archival research. Archival research relies on looking at past records or data sets to look for interesting patterns or relationships. For example, a researcher might access the academic records of all individuals who enrolled in college within the past ten years and...
17.1K
Applications of Integration to Probability Density Functions01:27

Applications of Integration to Probability Density Functions

Continuous probability distributions are used to model random variables that can take on any real value within a specified range. These variables do not take on isolated or countable values but rather exist on a continuum. For example, the height of an individual can be measured with increasing precision—such as 163.5 or 165.25 centimeters—demonstrating that height is a continuous random variable.The behavior of such variables is described using a probability density function (PDF),...
55
Probability Laws01:49

Probability Laws

Overview
44.3K
What Are Outliers?01:12

What Are Outliers?

Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
5.1K
Outliers and Influential Points01:08

Outliers and Influential Points

An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
6.3K
Ranks01:02

Ranks

Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
503