Related Experiment Videos
CellDuality: Unlocking Biological Reasoning in LLMs with Self-Supervised RLVR
Yuhang Chen1, Zhen Tan2, Ruichen Zhang1
1University of North Carolina at Chapel Hill.
Summary
CellDuality uses a novel self-supervised framework for large language models (LLMs) to achieve robust biological reasoning. This approach enables complex biological insights without needing verifiable outcomes, advancing computational biology.
Area of Science:
- Computational Biology
- Artificial Intelligence
- Genomics
Background:
- Developing generalist large language models (LLMs) for complex biological reasoning is a key challenge in computational biology.
- Existing LLMs struggle with open-ended, mechanistic reasoning, particularly in biology where outcomes are often non-verifiable.
- Reinforcement Learning from Verifiable Rewards (RLVR) shows promise but is limited by the infeasibility of verifying most biological outcomes.
Purpose of the Study:
- Introduce CellDuality, a self-supervised framework enabling LLM agents for robust reasoning in single-cell biology.
- Develop a method for training LLMs in biology using intrinsic rewards derived from a self-verification process.
- Enable scalable training of biological foundation models without requiring ground-truth verification labels.
Main Methods:
- CellDuality employs a complementary task duality principle, creating a bidirectional reasoning loop for self-verification.
- The framework involves a forward task (predicting biological outcomes) and an inverse task (reconstructing initial conditions from predictions).
- Intrinsic rewards from the reconstruction fidelity are used to align the LLM via reinforcement learning.
Main Results:
- CellDuality achieves state-of-the-art performance on diverse single-cell reasoning tasks, providing coherent biological explanations.
- The self-supervised approach significantly outperforms standard fine-tuning on out-of-distribution perturbation prediction.
- CellDuality narrows the performance gap to supervised RLVR baselines, demonstrating its effectiveness.
Conclusions:
- CellDuality offers a novel path toward scalable training of biological foundation models.
- The self-supervised, self-verification framework enhances LLM reasoning capabilities in biology.
- This work overcomes limitations of RLVR in biology by using intrinsic rewards from task duality.