Results 191 to 200 of about 71,442 (256)
Some of the next articles are maybe not open access.

On undiscounted semi-Markov decision processes with absorbing states

Mathematical Methods of Operations Research, 2016
zbMATH Open Web Interface contents unavailable due to conflicting licenses.
P. Mondal
semanticscholar   +2 more sources

Denumerable Undiscounted Semi-Markov Decision Processes with Unbounded Rewards

Mathematics of Operations Research, 1983
This paper establishes the existence of a solution to the optimality equations in undis-counted semi-Markov decision models with countable state space, under conditions generalizing the hitherto obtained results. In particular, we merely require the existence of a finite set of states in which every pair of states can reach each other via some ...
A. Federgruen, P. Schweitzer, H. Tijms
semanticscholar   +3 more sources

On Average Reward Semi-Markov Decision Processes with a General Multichain Structure

Mathematics of Operations Research, 2004
In this paper we investigate average reward semi-Markov decision processes with a general multichain structure using a data-transformation method. By solving the transformed discrete-time average Markov decision processes, we can obtain significant and interesting information on the original average semi-Markov decision processes. If the original semi-
Jianyong Liu, Xiaobo Zhao
exaly   +3 more sources

A policy gradient method for semi-Markov decision processes with application to call admission control

European Journal of Operational Research, 2007
Solving a semi-Markov decision process (SMDP) using value or policy iteration requires precise knowledge of the probabilistic model and suffers from the curse of dimensionality.
Arnaud Doucet
exaly   +2 more sources

An inverse reinforcement learning algorithm for semi-Markov decision processes

2017 IEEE Symposium Series on Computational Intelligence (SSCI), 2017
In this paper, we study the inverse reinforcement learning (IRL) algorithm for semi-Markov decision processes (SMDPs) with average reward based on the performance sensitivity analysis. By analyzing the structural form of the performance difference formula between any two different policies, we utilize the expert policy to transform the IRL problems of ...
C. Tan, Yanjie Li, Yuhu Cheng
semanticscholar   +2 more sources

Optimum Maintenance Policy Using Semi-Markov Decision Processes

Electric Power Systems Research, 2006
A method is presented to solve for the optimum maintenance policy of repairable power equipment. The approach uses a continuous-time semi-Markov process (SMP) to first find the optimal maintenance rate for maximum availability of the equipment. Then a semi-Markov decision process (SMDP) is utilized to determine whether maintenance should be performed ...
Curtis L. Tomasevicz, S. Asgarpoor
semanticscholar   +2 more sources

Average Reward Reinforcement Learning for Semi-Markov Decision Processes

International Conference on Neural Information Processing, 2017
In this paper, we study new reinforcement learning (RL) algorithms for Semi-Markov decision processes (SMDPs) with an average reward criterion. Based on the discrete-time type Bellman optimality equation, we use incremental value iteration (IVI), stochastic shortest path (SSP) value iteration and bisection algorithms to derive novel RL algorithms in a ...
Jiayuan Yang   +3 more
semanticscholar   +2 more sources

SEMI-MARKOV DECISION PROCESSES

Probability in the Engineering and Informational Sciences, 2007
Considered are semi-Markov decision processes (SMDPs) with finite state and action spaces. We study two criteria: the expected average reward per unit time subject to a sample path constraint on the average cost per unit time and the expected time-average variability.
M. Baykal-Gürsoy, K. Gürsoy
openaire   +2 more sources

Home - About - Disclaimer - Privacy