Related Experiment Video
Updated: Nov 13, 2025

A Tactile Automated Passive-Finger Stimulator TAPS
Published on: June 3, 2009
Learning parametric policies and transition probability models of markov decision processes from data.
Tingting Xu1, Henghui Zhu1, Ioannis Ch Paschalidis2
1Center for Information and Systems Engineering, Boston University, Boston, 02215, United States.
This study introduces new algorithms for learning Markov Decision Processes (MDPs) from data. The methods efficiently estimate policies and transition probabilities, achieving low regret with minimal samples.
Area of Science:
- Machine Learning
- Reinforcement Learning
- Statistical Modeling
Background:
- Markov Decision Processes (MDPs) are fundamental in sequential decision-making.
- Estimating MDP models from observational data is challenging.
- Parametric assumptions on sparse features are often necessary for tractability.
Purpose of the Study:
- To develop algorithms for estimating the policy and transition probability model of an MDP from state, action, next state tuples.
- To establish theoretical guarantees on the performance of the proposed estimation methods.
- To demonstrate the practical applicability of the algorithms through real-world examples.
Main Methods:
- Regularized maximum likelihood estimation for transition probability models.
- Regularized maximum likelihood estimation for policies.
- Theoretical analysis to establish an upper bound on regret.
- Sample complexity analysis to determine the number of samples required for low regret.
Main Results:
- Two novel regularized maximum likelihood estimation algorithms are proposed.
- An upper bound on the regret of the estimated policy is derived.
- A sample complexity result demonstrates that low regret can be achieved with limited data.
- The algorithms are validated on healthcare and robot navigation problems.
Conclusions:
- The proposed methods provide an effective approach for learning MDPs from data.
- The theoretical guarantees ensure efficient learning and good performance.
- The demonstrated applications highlight the practical utility of the algorithms in complex domains.
Related Concept Videos
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Probability Laws

