Results 151 to 160 of about 9,481,834 (176)
Some of the next articles are maybe not open access.
On the Optimality Conditions for Semi-Markov Decision Processes
1977The paper presents a recurrence formula for the difference between expected rewards and sojourn times generated by N transitions of a semi-Markov decision process with finite state space. Using the recurrence formula convergence of policy iteration method can be easily verified and also necessary and sufficient optimality conditions for average optimal
openaire +1 more source
Semi-Markov decision processes with a reachable state-subset
Optimization, 1989We consider the problem of minimizing the long-run average expected cost per unit time in a semi-Markov decision process with arbitrary state and action space, Assuming the existence .of a Borel subset of state space called a reachable state-subset, we derive the optimality equation for the unbounded costs.
openaire +1 more source
Deterministic policy gradient algorithms for semi‐Markov decision processes
International Journal of Intelligent Systems, 2021Ashkan Haji Hosseinloo +1 more
openaire +2 more sources
Learning Automaton for Finite Semi-Markov Decision Processes
1983A finite semi-Markov decision process is studied to maximize the expected average reward. The semi-Markov kernel of the process depends on an unknown parameter taking values in a subset [a, b] of ℝS. A controller modelled as a learning automaton updates sequentially the probabilities of generating decisions based on the observed decisions, states, and ...
openaire +1 more source
Reinforcement learning with options in semi Markov decision processes
2021Treball fi de màster de: Master in Intelligent Interactive ...
openaire +1 more source
Semi-Markov Decision-Making Processes with Vector Gains
Theory of Probability & Its Applications, 1984Vinogradskaya, T. M. +2 more
openaire +3 more sources
Cost Rate Heuristics for Semi-Markov Decision Processes
1991National Research ...
Glazebrook, K.D. +2 more
openaire +1 more source
A basic formula for performance gradient estimation of semi-Markov decision processes
European Journal of Operational Research, 2013Yanjie Li
exaly

