Related Experiment Video
Updated: Aug 6, 2026

Creating Objects and Object Categories for Studying Perception and Perceptual Learning
Published on: November 2, 2012
What topological and geometric structure do biological foundation models learn? Evidence from 141 hypotheses
1Department of Computer Science, University of Tübingen, Tübingen, Germany.
Abstract:
When biological foundation models like scGPT and Geneformer learn to process single-cell gene expression, what kind of geometric and topological structure forms in their internal representations? Is that structure biologically meaningful, or merely an artifact of training?. We address these questions through autonomous large-scale hypothesis screening: an AI-driven executor-brainstormer loop that proposed, tested, and refined 141 geometric and topological hypotheses across 52 iterations, covering persistent homology, manifold distances, cross-model alignment, community structure, directed topology, and more-all with explicit null controls and disjoint gene-pool splits. Three principal findings emerge. First, the models learn genuine geometric structure: gene embedding neighborhoods exhibit non-trivial topology (persistent homology significant in 11/12 transformer layers at p < 0.05 even in the weakest domain, and 12/12 in the other two), a multi-level distance hierarchy where manifold-aware metrics outperform Euclidean distance for identifying regulatory gene pairs, and graph-community partitions that track known transcription factor-target relationships. Second, this structure is shared across independently trained models: CCA alignment between scGPT and Geneformer yields canonical correlation of 0.80 and gene retrieval accuracy of 72%-yet no method among 19 tested could reliably recover gene-level correspondences, revealing that the models agree on the "shape" of gene space but not on precise gene placement. Third, the structure is more localized than it first appears: under the most stringent null controls (simultaneous auditing against all null families), robust signal concentrates in immune tissue, while lung and external-lung signals become fragile. These results-especially the carefully documented negatives among 141 hypotheses-calibrate what we can and cannot extract from biological model geometry, and demonstrate how autonomous screening can efficiently map the boundary between real structure and statistical artifact.
Related Concept Videos
The Evidence for Evolution
Structuralism
Titchener's approach to structuralism was unique. He employed introspection, a method...
The Tree of Life - Bacteria, Archaea, Eukaryotes
Evolutionary Relationships through Genome Comparisons
Null and Alternative Hypotheses
The null hypothesis, denoted by H0 is a statement of no difference between the variables—they are not related. This can often be considered the status quo. As a result if you cannot accept the null, it requires some action.
The alternative hypothesis, denoted by H1 or Ha, is a claim about the population that is...
DNA Topoisomerases
Types and Mechanism of action
Topoisomerases are divided into two main types. Type I...

