Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reinforcement Schedules01:24

Reinforcement Schedules

Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Lagrange Multipliers: Problem Solving01:30

Lagrange Multipliers: Problem Solving

A silo with a cylindrical base, flat bottom, and hemispherical roof is a common design in agricultural and industrial storage due to its structural efficiency and ease of construction. Optimizing its dimensions to maximize storage capacity for a given amount of material—i.e., a fixed surface area—is a classic problem in applied calculus and engineering design. The key parameters are the radius r of the base and the height h of the cylindrical section.The total volume of the silo is obtained by...
Observational Learning01:12

Observational Learning

Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
Reinforcement01:23

Reinforcement

Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Distributed Loads: Problem Solving01:21

Distributed Loads: Problem Solving

Beams are structural elements commonly employed in engineering applications requiring different load-carrying capacities. The first step in analyzing a beam under a distributed load is to simplify the problem by dividing the load into smaller regions, which allows one to consider each region separately and calculate the magnitude of the equivalent resultant load acting on each portion of the beam. The magnitude of the equivalent resultant load for each region can be determined by calculating...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Wavelet Probabilistic Neural Networks.

IEEE transactions on neural networks and learning systems·2022
Same author

Structural generative descriptions for time series classification.

IEEE transactions on cybernetics·2014
See all related articles

Related Experiment Videos

Reinforcement learning for resource allocation in LEO satellite networks.

Wipawee Usaha1, Javier A Barria

  • 1School of Telecommunication Engineering, Suranaree University of Technology, Nakorn Ratchasima 30000, Thailand. wipawee@sut.ac.th

IEEE Transactions on Systems, Man, and Cybernetics. Part B, Cybernetics : a Publication of the IEEE Systems, Man, and Cybernetics Society
|June 7, 2007
PubMed
Summary

We developed reinforcement learning algorithms for low Earth orbit (LEO) satellite networks to improve call routing. These methods significantly increase average revenue compared to existing techniques while reducing computational demands.

Related Experiment Videos

Area of Science:

  • Computer Science
  • Telecommunications Engineering
  • Operations Research

Background:

  • Low Earth orbit (LEO) satellite networks face challenges in efficient call admission and routing.
  • Traditional dynamic programming (DP) solutions for these problems are computationally prohibitive for large-scale systems.
  • Semi-Markov decision process (SMDP) formulations offer improved performance but require efficient solution methods.

Purpose of the Study:

  • To develop and evaluate online decision-making algorithms for call admission and routing in LEO satellite networks.
  • To overcome the computational limitations of dynamic programming (DP) for SMDP-based routing problems.
  • To enhance network performance metrics, including long-term average revenue and resource utilization.

Main Methods:

  • Development of two reinforcement learning (RL) algorithms: an actor-critic method with temporal-difference (TD) learning and an optimistic TD learning (critic-only) method.
  • Assessment of RL algorithms against conventional routing methods using numerical studies.
  • Evaluation of performance based on storage requirements, computational complexity, computational time, and average revenue function penalizing blocked calls.

Main Results:

  • Reinforcement learning (RL) algorithms significantly enhance performance compared to existing routing methods.
  • The proposed RL framework achieves up to 56% higher average revenue.
  • The algorithms demonstrate improved efficiency in terms of storage, computational complexity, and time.

Conclusions:

  • Reinforcement learning (RL) provides an effective and computationally feasible approach for call admission and routing in LEO satellite networks.
  • The developed RL algorithms offer a superior alternative to conventional methods, balancing performance gains with resource efficiency.
  • These findings pave the way for more optimized and profitable LEO satellite network operations.