Related Experiment Video
Updated: May 16, 2026

Working Memory Training for Older Participants: A Control Group Training Regimen and Initial Intellectual Functioning Assessment
Published on: September 20, 2020
Latent Chain-of-Thought for Visual Reasoning
Guohao Sun1,2, Hang Hua2,3, Jian Wang2
1Rochester Institute of Technology.
None:
Chain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such as SFT, PPO, and GRPO may not generalize well across unseen reasoning tasks and heavily rely on a biased reward model. To address this challenge, we reformulate reasoning in LVLMs as posterior inference and propose a scalable training algorithm based on amortized variational inference. By leveraging diversity-seeking reinforcement learning algorithms, we introduce a novel sparse reward function for token-level learning signals that encourage diverse, high-likelihood latent CoT, overcoming deterministic sampling limitations and avoiding reward hacking. Additionally, we implement a Bayesian inference-scaling strategy that replaces costly Best-of-N and Beam Search with a marginal likelihood to efficiently rank optimal rationales and answers. We empirically demonstrate that the proposed method enhances the state-of-the-art LVLMs on seven reasoning benchmarks, in terms of effectiveness, generalization, and interpretability. The code is available at https://github.com/heliossun/LaCoT.
Related Concept Videos
Inductive Reasoning
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Reasoning
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Reason and Intuition
Deductive Reasoning
For example, a researcher can deduce specific predictions...
Counterfactual Thinking
Visual Agnosia

