Related Experiment Video
Updated: Aug 6, 2026

Creating Objects and Object Categories for Studying Perception and Perceptual Learning
Published on: November 2, 2012
What topological and geometric structure do biological foundation models learn? Evidence from 141 hypotheses
1Department of Computer Science, University of Tübingen, Tübingen, Germany.
Biological foundation models learn meaningful geometric and topological structures in gene expression data, revealing shared patterns across models. However, the precise gene placement within these structures remains ambiguous, with robust signals concentrated in immune tissues.
Area of Science:
- Computational Biology
- Single-cell Genomics
- Machine Learning
Background:
- Biological foundation models (e.g., scGPT, Geneformer) process single-cell gene expression data.
- Understanding the internal geometric and topological structures learned by these models is crucial.
- Distinguishing biologically meaningful structure from training artifacts is a key challenge.
Purpose of the Study:
- To investigate the geometric and topological structures within biological foundation models.
- To determine if these learned structures are biologically meaningful or artifacts of training.
- To employ autonomous hypothesis screening for large-scale investigation.
Main Methods:
- An AI-driven executor-brainstormer loop screened 141 geometric and topological hypotheses.
- Methods included persistent homology, manifold distances, cross-model alignment, and community structure analysis.
- Explicit null controls and disjoint gene-pool splits were utilized for rigorous validation.
Main Results:
- Models learn genuine geometric structure, with non-trivial topology in gene embeddings and a distance hierarchy.
- Manifold-aware metrics outperformed Euclidean distance for identifying regulatory gene pairs.
- Structure is shared across independently trained models (scGPT, Geneformer) with high alignment, but precise gene placement differs.
- Robust topological signal is concentrated in immune tissues, while lung tissue signals are fragile under stringent controls.
Conclusions:
- Biological foundation models capture genuine, shared geometric and topological structure in gene expression data.
- The precise localization of genes within this learned space is less consistent and more context-dependent.
- Autonomous screening effectively differentiates real biological structure from statistical artifacts in model representations.
Related Concept Videos
The Evidence for Evolution
Structuralism
Titchener's approach to structuralism was unique. He employed introspection, a method...
The Tree of Life - Bacteria, Archaea, Eukaryotes
Evolutionary Relationships through Genome Comparisons
Null and Alternative Hypotheses
The null hypothesis, denoted by H0 is a statement of no difference between the variables—they are not related. This can often be considered the status quo. As a result if you cannot accept the null, it requires some action.
The alternative hypothesis, denoted by H1 or Ha, is a claim about the population that is...
DNA Topoisomerases
Types and Mechanism of action
Topoisomerases are divided into two main types. Type I...

