Related Experiment Videos
Application-driven pedagogical knowledge optimization of open-source LLMs via reinforcement learning and supervised
Navan Preet Singh1, Xiaokun Wang2, Anurag Garikipati3
1Department of Research and Development, Forta, Houston, TX, United States.
Abstract:
We present a multi-stage optimization strategy combining reinforcement learning (RL) and supervised fine-tuning (SFT) to enhance the pedagogical knowledge of Large Language Models (LLMs), thereby providing a technically grounded example of how open-source pedagogical LLMs can be optimized for deployment in diverse educational settings, including institutions with constrained resources. Our approach produces EduQwen 32B-RL1, EduQwen 32B-SFT, and EduQwen 32B-SFT-RL2: (1) first-stage RL optimization implementing progressive difficulty training, focusing on challenging examples, and employing extended reasoning rollouts to facilitate adaptive scaffolding, prioritizing pedagogical steering over direct answer provision; (2) SFT that leverages the RL-trained model to synthesize high-quality training data with difficulty-weighted sampling; and (3) optional second-stage RL refinement. This application-driven family of open-source pedagogical LLMs, built on a dense Qwen3-32B backbone, achieves 96.52% accuracy on the Cross-Domain Pedagogical Knowledge (CDPK) Benchmark, establishing new state-of-the-art (SOTA) performance on the CDPK subset of the interactive Pedagogy Benchmark Leaderboard as of March 2026, with a 5.97 percentage points accuracy gain over the then-reported Gemini-3 Pro's score (previous leader with 90.55% accuracy), under the respective documented evaluation protocols. Critically, our 32-billion-parameter models demonstrate that domain-specialized optimization of mid-sized open-source LLMs can outperform much larger general-purpose systems on pedagogical knowledge benchmarks, indicating a scalable technical approach for wider AI deployment and potential pedagogical support, while preserving the transparency, customizability, and cost-efficiency required for responsible educational AI development, with deployment effectiveness depending on teacher judgment, learner context, and authentic instructional integration across diverse global learning environments.
Related Concept Videos
Observational Learning
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Purposive Learning
Introduction to Learning
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
Associative Learning
Classical conditioning, also known...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example: