Related Experiment Videos
Intelligent optimization of personalized learning path based on transformer and reinforcement learning.
1School of Information and Intelligent Engineering, Zhangjiajie College, Zhangjiajie, 427000, China.
Scientific Reports
|July 14, 2026
Summary
A new Transformer and Reinforcement Learning model (TFRL-Path) enhances personalized learning paths by dynamically adapting to learner progress. This improves recommendation accuracy and learning outcomes, outperforming existing methods.
Area of Science:
- Artificial Intelligence in Education
- Machine Learning for Educational Data Mining
- Educational Technology
Background:
- Smart education platforms generate vast learning data, but static recommendations fail to address individual learner differences.
- Existing learning path methods lack dynamic state awareness and long-term path optimization capabilities.
- Personalized learning requires adaptive strategies that consider evolving learner states and cognitive progress.
Purpose of the Study:
- To develop an optimized personalized learning path (PLP) model using Transformer and Reinforcement Learning (RL).
- To enhance learning resource recommendation accuracy, path continuity, and learning completion rates.
- To generate learning sequences aligned with individual cognitive processes via dynamic state awareness and educational constraints.
Main Methods:
- Unified framework encoding learning behaviors, knowledge, feedback, resource access, and time intervals.
- Transformer model for capturing long-range dependencies and dynamic state features.
- Reinforcement Learning for sequential policy decision-making and optimization within educational constraints (prerequisites, difficulty, load).
Main Results:
- TFRL-Path model demonstrated superior performance across three public datasets (EdNet, ASSISTments, Junyi).
- Achieved a hit ratio of 0.846 on EdNet, outperforming UniKT by 0.025.
- Showcased improved mean average precision (0.737 on ASSISTments), completion rate (0.801 on ASSISTments), prerequisite satisfaction (0.872), and reduced repetition (0.121).
Conclusions:
- The integrated Transformer and Reinforcement Learning approach significantly improves personalized learning path recommendation quality.
- Dynamic state awareness, sequential policy optimization, and educational constraints are crucial for effective adaptive learning systems.
- TFRL-Path offers a promising solution for creating more effective and engaging personalized learning experiences.
Related Concept Videos
Introduction to Learning
Learning is the process of acquiring knowledge or skills through practice or experience, leading to long-lasting behavioral changes. This acquisition occurs through interaction with the environment and requires practice or experience. For instance, mastering a skill such as surfing requires considerable practice and experience, highlighting the essential role of repeated interactions with the environment in learning.
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
Cognitive Learning
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Observational Learning
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
Purposive Learning
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a bonus...
The Ideal Transformer
In single-phase two-winding transformers, two windings are coiled around a magnetic core characterized by cross-sectional area A and magnetic permeability μ. A phasor current i1 enters the left winding while i2 exits the right winding, establishing the fundamental working of the transformer through electromagnetic principles.
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's tangential component...
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's tangential component...
Transformers
A device that transforms voltages from one value to another using induction is called a transformer. A transformer consists of two separate coils, or windings, wrapped around the same soft iron core. However, they are electrically insulated from each other.
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...