Related Experiment Video
Updated: May 6, 2026

Study Motor Skill Learning by Single-pellet Reaching Tasks in Mice
Published on: March 4, 2014
Cell-o1 : training LLMs to solve single-cell reasoning puzzles with reinforcement learning
Yin Fang1, Qiao Jin1, Guangzhi Xiong2
1Division of Intramural Research, National Library of Medicine, National Institutes of Health, 8600 Rockville Pike, Bethesda, MD 20894, United States.
Motivation:
Large language models (LLMs) have demonstrated strong general reasoning abilities, but applying them to domain-specific tasks such as analysing single-cell RNA sequencing data remains a challenge. A central task in this domain is cell type annotation, which is critical for understanding cellular heterogeneity. Although recent foundation models attempt to automate this process, they typically annotate cells independently, without considering batch-level context or providing explanatory reasoning. To address this limitation, we introduce the CellPuzzles benchmark, which reformulates cell type annotation as a batch-level reasoning task. CellPuzzles spans diverse tissues, diseases, and donor conditions, and requires reasoning across the batch-level cellular context to ensure label uniqueness.
Results:
We find that off-the-shelf LLMs struggle on this task, with the best baseline (OpenAI o1) achieving only 19.0% batch-level accuracy. To fill this gap, we propose Cell-o1, a 7B LLM trained via supervised fine-tuning on distilled reasoning traces, followed by reinforcement learning with batch-level rewards. Cell-o1 achieves state-of-the-art performance, outperforming OpenAI o1 by over 73% and generalizing well across contexts. Further analysis of training dynamics and reasoning behaviors provides insights into batch-level annotation performance and emergent expert-like reasoning.
Availability And Implementation:
Code and data are available at https://github.com/ncbi-nlp/cell-o1.
Related Concept Videos
Machines: Problem Solving II
Observational Learning
Machines: Problem Solving I
The toggle clamp system is a machine structure consisting of movable, pin-connected multi-force members that form a stabilized system to transmit forces. The...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Associative Learning
Classical conditioning, also known...
Purposive Learning

