Wrangling distributed computing for high-throughput environmental science: An introduction to HTCondor

Richard A Erickson1, Michael N Fienen2, S Grace McCalla1

  • 1Upper Midwest Environmental Sciences Center, United States Geological Survey, La Crosse, Wisconsin, United States of America.

Summary

High-throughput computing (HTC) offers a scalable solution for computationally intensive research in biology and environmental science. This approach distributes tasks across multiple computers, overcoming limitations of traditional methods.

Related Concept Videos

Introduction to Normal Distributions01:29

Introduction to Normal Distributions

Standardized test scores often follow a symmetric distribution that can be modeled with the normal distribution, a fundamental concept in statistics. This distribution is particularly useful for interpreting test performance fairly across populations, as it provides a mathematical framework for understanding variability and central tendency in large datasets.From Histogram to Frequency DistributionRaw test data are often displayed using histograms, where the height of each bar represents the...
75
Psychology as a Science01:13

Psychology as a Science

Psychology, as a scientific discipline, aims to understand the mind and behavior through rigorous and systematic methods. The foundation of psychological research is evidence-based, relying heavily on the scientific method to derive and validate knowledge. This structured approach ensures that findings are reliable, valid, and applicable to broader contexts.
The scientific method in psychology involves six critical steps: making observations, formulating hypotheses, conducting tests, analyzing...
3.9K
Overview of Biostatistics in Health Sciences01:19

Overview of Biostatistics in Health Sciences

Biostatistics involves the application of statistical techniques to scientific research in health-related fields, including biology and public health. These techniques are essential for designing studies, collecting data, and analyzing it to draw meaningful conclusions. Given the complexity of biological processes, particularly in studies involving human subjects, biostatistical methods are crucial for effectively organizing and interpreting data that might otherwise obscure underlying patterns...
5.3K
Statistical Package for the Social Sciences (SPSS)01:22

Statistical Package for the Social Sciences (SPSS)

The Statistical Package for the Social Sciences, or SPSS, is a data management and analysis software suite. Developed by SPSS Inc. in 1968 and acquired by IBM in 2009, this tool was initially designed for social science data analysis, evolving to serve a wider range of disciplines. It was later renamed to Statistical Product and Service Solutions.
SPSS streamlines the process from data preparation to analysis and reporting. It is characterized by its user-friendly interface, which conceals...
1.3K
Drug Distribution: Volume of Distribution01:25

Drug Distribution: Volume of Distribution

The volume of distribution refers to the theoretical volume necessary to contain the entire amount of an administered drug at the same concentration observed in the blood plasma. The body's intracellular fluid compartment, which makes up two-thirds of the total body water, is contrasted with the extracellular fluid compartment—comprising plasma and interstitial fluid—that accounts for one-third. The volume of distribution can vary depending on the characteristics of the drug.
7.5K
Variation: Normal Distribution, Range, and Standard Deviation02:32

Variation: Normal Distribution, Range, and Standard Deviation

In the field of psychology, there are several ways to organize measurements of a trait, feature, or characteristic (i.e., variables). Qualitative data, such as ethnicity, can be tabulated into a frequency count to provide information about the proportion, as well as the variety of groups in a sample or population. On the other hand, researchers can perform a wider set of calculations on quantitative data. The mean, mode, and median, for instance, are central tendency measures to identify a...
28.3K