Related Experiment Video
Updated: Apr 30, 2026

05:30
Large Scale Energy Efficient Sensor Network Routing Using a Quantum Processor Unit
Published on: September 8, 2023
1.3K
Fidelity-based probabilistic Q-learning for control of quantum systems.
Summary
A new fidelity-based probabilistic Q-learning (FPQL) method balances exploration and exploitation in reinforcement learning. This approach enhances control of quantum systems by avoiding local optima and speeding up learning.
Area of Science:
- Quantum Control
- Reinforcement Learning
- Machine Learning
Background:
- The exploration-exploitation dilemma is a central challenge in reinforcement learning, particularly for Q-learning algorithms.
- Effective strategies are needed to balance exploring new actions and exploiting known ones for optimal policy learning.
Purpose of the Study:
- To introduce a novel fidelity-based probabilistic Q-learning (FPQL) approach for reinforcement learning.
- To apply FPQL for the learning control of quantum systems.
- To address the exploration-exploitation balance in reinforcement learning for quantum applications.
Main Methods:
- Developed a probabilistic Q-learning (PQL) algorithm to illustrate probabilistic action selection.
- Introduced the fidelity-based probabilistic Q-learning (FPQL) algorithm, using fidelity to guide learning.
- Iteratively updated action selection probabilities based on learning progress.
Main Results:
- FPQL demonstrated a superior balance between exploration and exploitation compared to traditional methods.
- The algorithm successfully avoided local optimal policies in quantum system control.
- FPQL accelerated the learning process for controlling quantum systems.
Conclusions:
- The FPQL approach offers a natural and effective solution to the exploration-exploitation problem in reinforcement learning.
- FPQL shows significant promise for advancing the field of quantum control through improved learning strategies.
- The method's ability to avoid local optima and speed up learning makes it a valuable tool for complex control tasks.
Related Concept Videos
Propagation of Uncertainty from Systematic Error
1.4K
The atomic mass of an element varies due to the relative ratio of its isotopes. A sample's relative proportion of oxygen isotopes influences its average atomic mass. For instance, if we were to measure the atomic mass of oxygen from a sample, the mass would be a weighted average of the isotopic masses of oxygen in that sample. Since a single sample is not likely to perfectly reflect the true atomic mass of oxygen for all the molecules of oxygen on Earth, the mass we obtain from this...
1.4K
Propagation of Uncertainty from Random Error
1.9K
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
1.9K
The Quantum-Mechanical Model of an Atom
47.1K
Shortly after de Broglie published his ideas that the electron in a hydrogen atom could be better thought of as being a circular standing wave instead of a particle moving in quantized circular orbits, Erwin Schrödinger extended de Broglie’s work by deriving what is now known as the Schrödinger equation. When Schrödinger applied his equation to hydrogen-like atoms, he was able to reproduce Bohr’s expression for the energy and, thus, the Rydberg formula governing...
47.1K
Observational Learning
1.5K
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.5K
BIBO stability of continuous and discrete -time systems
1.1K
System stability is a fundamental concept in signal processing, often assessed using convolution. For a system to be considered bounded-input bounded-output (BIBO) stable, any bounded input signal must produce a bounded output signal. A bounded input signal is one where the modulus does not exceed a certain constant at any point in time.
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
1.1K
Conservation of Energy in Control Volume
1.1K
Consider a turbine operating under steady-flow conditions. The control volume is drawn around the turbine, with fluid entering at one point and exiting at another. The turbine extracts energy from the fluid, which performs mechanical work (shaft work).
For steady flow systems, the time derivative of the stored energy becomes zero since there is no energy accumulation within the control volume. This simplifies the energy equation to:
For steady flow systems, the time derivative of the stored energy becomes zero since there is no energy accumulation within the control volume. This simplifies the energy equation to:
1.1K
