Related Experiment Video
Updated: Apr 10, 2026

Rup (RNA-seq Usability Assessment Pipeline) - Quality Control for Bulk RNA-seq Experiments in Eukaryotes
Published on: November 7, 2025
Zobrist hash-based duplicate detection in symbolic regression
1Software Engineering, University of Applied Sciences Upper Austria School of Informatics Communications and Media, Hagenberg, Hagenberg, Austria.
Abstract:
Symbolic regression encompasses a family of search algorithms that aim to discover the best-fitting function for a set of data without requiring an a priori specification of the model structure. The most successful and commonly used technique for symbolic regression is genetic programming (GP), an evolutionary search method that evolves a population of mathematical expressions through the mechanism of natural selection. In this work, we analyse the efficiency of the evolutionary search in GP and show that many points in the search space are revisited and re-evaluated multiple times by the algorithm, leading to wasted computational effort. We address this issue by introducing a caching mechanism based on the Zobrist hash, a type of hashing frequently used in abstract board games for the efficient construction and subsequent update of transposition tables. We implement our caching approach using the open-source framework Operon and demonstrate its performance on a selection of real-world regression problems, where we observe up to 34% speedups without any detrimental effects on search quality. The hashing approach represents a straightforward way of improving runtime performance while also offering some interesting possibilities for adjusting the search strategy based on cached information. This article is part of the discussion meeting issue 'Symbolic regression in the physical sciences'.
More Related Videos
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Related Concept Videos
Quantifying and Rejecting Outliers: The Grubbs Test
Wald-Wolfowitz Runs Test II
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
Routh-Hurwitz Criterion II
The first scenario occurs when a singular zero appears in the first column of the Routh table. This situation creates a division by zero issues. To resolve this, a small positive or negative number, denoted as epsilon (∈), is substituted for the zero. The stability analysis proceeds by assuming a sign for ∈. If ∈ is positive, any sign change in the first...
Hypothesis Test for Test of Independence
H0: The two variables (factors)...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...