Results 171 to 180 of about 38,584 (203)
Some of the next articles are maybe not open access.

An inverse reinforcement learning algorithm for semi-Markov decision processes

2017 IEEE Symposium Series on Computational Intelligence (SSCI), 2017
In this paper, we study the inverse reinforcement learning (IRL) algorithm for semi-Markov decision processes (SMDPs) with average reward based on the performance sensitivity analysis. By analyzing the structural form of the performance difference formula between any two different policies, we utilize the expert policy to transform the IRL problems of ...
Chuanfang Tan, Yanjie Li, Yuhu Cheng
openaire   +1 more source

On Average Reward Semi-Markov Decision Processes with a General Multichain Structure

Mathematics of Operations Research, 2004
In this paper we investigate average reward semi-Markov decision processes with a general multichain structure using a data-transformation method. By solving the transformed discrete-time average Markov decision processes, we can obtain significant and interesting information on the original average semi-Markov decision processes. If the original semi-
Jianyong Liu, Xiaobo Zhao
openaire   +2 more sources

Performance Optimization of Semi-Markov Decision Processes with Discounted-cost Criteria

European Journal of Control, 2008
zbMATH Open Web Interface contents unavailable due to conflicting licenses.
Baoqun Yin   +3 more
openaire   +1 more source

Semi-Markov decision processes with polynomial reward

Journal of Applied Probability, 1982
A semi-Markov decision process, with a denumerable multidimensional state space, is considered. At any given state only a finite number of actions can be taken to control the process. The immediate reward earned in one transition period is merely assumed to be bounded by a polynomial and a bound is imposed on a weighted moment of the next state reached
openaire   +1 more source

Finite horizon semi-Markov decision processes with application to maintenance systems

European Journal of Operational Research, 2011
zbMATH Open Web Interface contents unavailable due to conflicting licenses.
Yonghui Huang, Xianping Guo
openaire   +1 more source

Average Reward Reinforcement Learning for Semi-Markov Decision Processes

2017
In this paper, we study new reinforcement learning (RL) algorithms for Semi-Markov decision processes (SMDPs) with an average reward criterion. Based on the discrete-time type Bellman optimality equation, we use incremental value iteration (IVI), stochastic shortest path (SSP) value iteration and bisection algorithms to derive novel RL algorithms in a ...
Jiayuan Yang   +3 more
openaire   +1 more source

Reinforcement learning for semi-Markov decision processes with applications

2023
This thesis focuses on semi-Markov decision processes and their connection with Reinforcement Learning via Q-learning technique. We start by discussing some general ideas around Machine Learning, Reinforcement Learning and Hierarchical Reinforcement Learning.
openaire   +1 more source

Constrained Discounted Semi-Markov Decision Processes

2002
This paper reduces problems on the existence and the finding of optimal policies for multiple criterion discounted SMDPs to similar problems for MDPs. We prove this reduction and illustrate it by extending to SMDPs several results for constrained discounted MDPs.
openaire   +1 more source

Semi-Markov decision processes with a reachable state-subset

Optimization, 1989
We consider the problem of minimizing the long-run average expected cost per unit time in a semi-Markov decision process with arbitrary state and action space, Assuming the existence .of a Borel subset of state space called a reachable state-subset, we derive the optimality equation for the unbounded costs.
openaire   +1 more source

On the Optimality Conditions for Semi-Markov Decision Processes

1977
The paper presents a recurrence formula for the difference between expected rewards and sojourn times generated by N transitions of a semi-Markov decision process with finite state space. Using the recurrence formula convergence of policy iteration method can be easily verified and also necessary and sufficient optimality conditions for average optimal
openaire   +1 more source

Home - About - Disclaimer - Privacy