Related Experiment Video
Updated: Jun 25, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Working with benchmark datasets in the Cuby framework
Jan Řezáč1, Outi Vilhelmiina Kontkanen1, Martin Nováček1
1Institute of Organic Chemistry and Biochemistry, Czech Academy of Sciences, 160 00 Prague, Czech Republic.
Abstract:
The development and benchmarking of computational chemistry methods rely on comparison with benchmark data. More and larger benchmark datasets are becoming available, and working efficiently with them is a necessity. The Cuby framework provides rich functionality for working with datasets, comes with many ready-to-use predefined benchmark sets, and interfaces with a wide range of computational chemistry software packages. Here, we review the tools Cuby provides for working with datasets and provide examples of more advanced workflows, such as handling large numbers of computations on high performance computing resources and reusing previously computed data. Cuby has also been extended recently to include two important benchmark databases, NCIAtlas and GMTKN55.
Related Concept Videos
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Quartile
1; 1; 2; 2; 4; 6; 6.8; 7.2; 8; 8.3; 9; 10; 10; 11.5
The median or second quartile is seven. The lower half of the...
Data Collection II
Data Collection I
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Histogram
A histogram graph consists of contiguous (adjoining) boxes. The heights of the bars correspond to frequency values. The graph will have the same shape with respective labels. The...

