Inference Strategies for Solving Semi-Markov Decision Processes
Semi-Markov decision processes are used to formulate many control problems and also play a key role in hierarchical reinforcement learning. In this chapter we show how to translate the decision making problem into a form that can instead be solved by inference and learning techniques.
Hoffman, M, de Freitas, N
core +4 more sources
Mixed Markov Decision Processes in a Semi-Markov Environment with Discounted Criterion [PDF]
zbMATH Open Web Interface contents unavailable due to conflicting licenses.
Hu, Qiying, Wang, Jinling
openaire +2 more sources
Hierarchical dialogue optimization using semi-Markov decision processes [PDF]
This paper addresses the problem of dialogue optimization on large search spaces. For such a purpose, in this paper we propose to learn dialogue strategies using multiple Semi-Markov Decision Processes and hierarchical reinforcement learning. This approach factorizes state variables and actions in order to learn a hierarchy of policies. Our experiments
Cuayáhuitl, Heriberto +3 more
openaire +5 more sources
The exponential cost optimality for finite horizon semi-Markov decision processes [PDF]
summary:This paper considers an exponential cost optimality problem for finite horizon semi-Markov decision processes (SMDPs). The objective is to calculate an optimal policy with minimal exponential costs over the full set of policies in a finite ...
Wen, Xian, Huo, Haifeng
core +1 more source
Risk probability optimization problem for finite horizon continuous time Markov decision processes with loss rate [PDF]
summary:This paper presents a study the risk probability optimality for finite horizon continuous-time Markov decision process with loss rate and unbounded transition rates.
Wen, Xian, Huo, Haifeng
core +1 more source
Mean-variance optimality for semi-Markov decision processes under first passage criteria [PDF]
summary:This paper deals with a first passage mean-variance problem for semi-Markov decision processes in Borel spaces. The goal is to minimize the variance of a total discounted reward up to the system's first entry to some target set, where the ...
Xiangxiang Huang +3 more
core +1 more source
A Fast-Pivoting Algorithm for Whittle’s Restless Bandit Index
The Whittle index for restless bandits (two-action semi-Markov decision processes) provides an intuitively appealing optimal policy for controlling a single generic project that can be active (engaged) or passive (rested) at each decision epoch, and ...
José Niño-Mora
doaj +1 more source
First passage risk probability optimality for continuous time Markov decision processes [PDF]
summary:In this paper, we study continuous time Markov decision processes (CTMDPs) with a denumerable state space, a Borel action space, unbounded transition rates and nonnegative reward function.
Wen, Xian, Huo, Haifeng
core +1 more source
Optimal Intervention in Semi-Markov-Based Asynchronous Probabilistic Boolean Networks
Synchronous probabilistic Boolean networks (PBNs) and generalized asynchronous PBNs have received significant attention over the past decade as a tool for modeling complex genetic regulatory networks.
Qiuli Liu +3 more
doaj +1 more source
Nonstationary Continuous Time Markov Decision Processes in a Semi-Markov Environment with Discounted Criterion [PDF]
This paper deals with the nonstationary continuous time Markov decision process in a semi-Markov environment with discounted criterion. The model can describe a system that itself can be modeled by a countable state nonstationary continuous time Markov ...
Hu, Q.Y.
core +1 more source

