Related Experiment Video
Updated: May 6, 2026

Study Motor Skill Learning by Single-pellet Reaching Tasks in Mice
Published on: March 4, 2014
Cell-o1 : training LLMs to solve single-cell reasoning puzzles with reinforcement learning
Yin Fang1, Qiao Jin1, Guangzhi Xiong2
1Division of Intramural Research, National Library of Medicine, National Institutes of Health, 8600 Rockville Pike, Bethesda, MD 20894, United States.
We developed Cell-o1, a large language model (LLM) that excels at cell type annotation for single-cell RNA sequencing data. Cell-o1 significantly improves accuracy by considering batch-level context, outperforming existing methods.
Area of Science:
- Computational Biology
- Genomics
- Artificial Intelligence
Background:
- Large language models (LLMs) show general reasoning but struggle with specialized tasks like single-cell RNA sequencing (scRNA-seq) data analysis.
- Cell type annotation is crucial for understanding cellular heterogeneity in scRNA-seq data.
- Current automated methods often lack batch-level context and explanatory reasoning.
Purpose of the Study:
- Introduce the CellPuzzles benchmark for batch-level cell type annotation.
- Develop a novel LLM, Cell-o1, to address limitations in current cell type annotation methods.
- Improve the accuracy and contextual understanding of cell type annotation in scRNA-seq data.
Main Methods:
- Reformulated cell type annotation as a batch-level reasoning task using the CellPuzzles benchmark.
- Developed Cell-o1, a 7B parameter LLM.
- Employed supervised fine-tuning with distilled reasoning traces and reinforcement learning with batch-level rewards for training Cell-o1.
Main Results:
- Off-the-shelf LLMs achieved low batch-level accuracy (19.0% for OpenAI o1).
- Cell-o1 significantly outperformed existing baselines, achieving over 73% improvement compared to OpenAI o1.
- Cell-o1 demonstrated strong generalization across diverse tissues, diseases, and donor conditions.
Conclusions:
- Cell-o1 represents a state-of-the-art approach for batch-aware cell type annotation in scRNA-seq data.
- The CellPuzzles benchmark facilitates the development and evaluation of LLMs for this task.
- Further analysis offers insights into LLM reasoning for biological data interpretation.
Related Concept Videos
Machines: Problem Solving II
Observational Learning
Machines: Problem Solving I
The toggle clamp system is a machine structure consisting of movable, pin-connected multi-force members that form a stabilized system to transmit forces. The...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Associative Learning
Classical conditioning, also known...
Purposive Learning

