Data quantity is more important than its spatial bias for predictive species distribution modelling
Willson Gaul1, Dinara Sadykova2, Hannah J White1
1School of Biology and Environmental Science, Earth Institute, University College Dublin, Dublin, Ireland.
Peerj
|December 14, 2020
Summary
Spatial bias in species distribution models (SDMs) can decrease prediction performance. However, sample size and the choice of modeling method are more critical factors than spatial bias for accurate SDM predictions.
Area of Science:
- Ecology
- Biodiversity Science
- Computational Biology
Background:
- Biological records are crucial for training species distribution models (SDMs).
- Spatial sampling bias is a common issue in biological data, potentially affecting SDM performance.
- Understanding the impact of bias on SDM predictions is essential for ecological research.
Purpose of the Study:
- To evaluate the impact of spatial sampling bias, sample size, and modeling method on species distribution model (SDM) prediction performance.
- To simulate biological recording processes with real-world spatial biases.
- To quantify the relative importance of these factors in determining SDM accuracy.
Main Methods:
- Simulated presence and absence data for virtual species.
- Incorporated realistic spatial sampling biases into simulated data.
- Assessed prediction performance across different sample sizes and various SDM methods.
Main Results:
- Spatial bias in training data was found to decrease SDM prediction performance.
- Sample size and the selection of the SDM method had a greater impact on prediction performance than spatial bias.
- The study highlights the interplay between data quality and modeling choices.
Conclusions:
- While spatial bias affects SDM performance, it is not the sole determinant of accuracy.
- Optimizing sample size and selecting appropriate modeling techniques are critical for robust species distribution modeling.
- Future research should consider both data biases and methodological choices for reliable ecological predictions.
Related Concept Videos
Distribution and Dispersion
23.6K
To understand intra-specific interactions in populations, scientists measure the spatial arrangement of species individuals. This geographic arrangement is known as the species distribution or dispersion. Highly territorial species exhibit a uniform distribution pattern, in which individuals are spaced at relatively equal distances from one another. Species that are highly tied to particular resources, such as food or shelter, tend to concentrate around those resources, and thus exhibit a...
23.6K
Distributions to Estimate Population Parameter
4.8K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.8K
Selected Data About Geographic Locations
149
Geographic Information Systems (GIS) rely on two core types of data: spatial data and attribute data.Spatial DataSpatial data defines the physical location of features within a coordinate system, typically expressed in terms of latitude and longitude. It provides precise positioning for elements like roads, rivers, or buildings.Attribute DataAttribute data complements spatial data by adding descriptive information about these features. For example, a road's spatial data includes its start and...
149
Choosing Between z and t Distribution
3.4K
The z and the Student t distribution estimate the population mean using the sample mean and standard deviation. However, to decide which distribution to use for a calculation, one needs to determine the sample size, the nature of the distribution, and whether the population standard deviation is known. If the population standard deviation is known and the population is normally distributed, or if the sample size is greater than 30, the z distribution is preferred. The Student t distribution is...
3.4K
Data: Types and Distribution
1.1K
In biostatistics, data are the observations collected for analysis. There are two main types: parametric and non-parametric. Parametric data, which include continuous (e.g., weight) and discrete numerical data (e.g., number of tablets), assume a particular distribution pattern, often the normal distribution. Non-parametric data do not adhere to a specific distribution and typically comprise nominal (e.g., gender) and ordinal categorical data (e.g., pain scale ratings).
Distributions in...
Distributions in...
1.1K
Habitat Fragmentation
20.4K
Habitat fragmentation describes the division of a more extensive, continuous habitat into smaller, discontinuous areas. Human activities such as land conversion, as well as slower geological processes leading to changes in the physical environment, are the two leading causes of habitat fragmentation. The fragmentation process typically follows the same steps: perforation, dissection, fragmentation, shrinkage, and attrition.
20.4K


