Integrated forecasting and deep reinforcement learning for price-based self-scheduling of PV-BESS: Utility-scale

Juan Pérez1, Gustavo Lobos1, Milena Bonacic1

  • 1Facultad de Ingeniería y Ciencias Aplicadas, Universidad de Los Andes, Santiago, Chile.

Plos One
|January 9, 2026
PubMed
Summary

This study validates Deep Reinforcement Learning (DRL) for battery energy storage systems (BESS) and photovoltaic (PV) plant control using real-world data. DRL agents significantly boosted profits and demonstrated adaptive operations, proving practical viability.

Related Concept Videos

Energy Budgets00:51

Energy Budgets

Organisms must balance energy intake with the energy required for growth, maintenance and reproduction. These trade-offs result in a variety of survivorship and reproductive strategies, including semelparity and iteroparity. Semelparous species, like annual plants, have only one reproductive episode in their lifetimes and consequently have short lifespans. Iteroparous species, by contrast, have many reproductive events during their lifetimes but have relatively few offspring. These two...
10.5K
Prediction Intervals01:03

Prediction Intervals

The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y. 
3.2K
Maxwell-Boltzmann Distribution: Problem Solving01:20

Maxwell-Boltzmann Distribution: Problem Solving

Individual molecules in a gas move in random directions, but a gas containing numerous molecules has a predictable distribution of molecular speeds, which is known as the Maxwell-Boltzmann distribution, f(v).
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
2.8K
Maximum Power Flow and Line Loadability01:23

Maximum Power Flow and Line Loadability

The maximum power flow for lossy transmission lines is derived using ABCD parameters in phasor form. These parameters create a matrix relationship between the sending-end and receiving-end voltages and currents, allowing the determination of the receiving-end current. This relationship facilitates calculating the complex power delivered to the receiving end, from which real and reactive power components are derived.
582
Reinforcement Schedules01:24

Reinforcement Schedules

Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
447
Applications of Integration to Find Consumer Surplus01:29

Applications of Integration to Find Consumer Surplus

In microeconomics, consumer surplus represents the economic gain that consumers experience when they purchase a good or service for less than the highest price they are willing to pay. This surplus arises from the characteristics of the demand function, which links the quantity of a good to the price consumers are willing to pay. As the quantity of a good increases, the price that consumers are willing to pay for each additional unit typically decreases, resulting in a downward-sloping demand...
2