A Markov Reward Process-Based Approach to Spatial Interpolation
The interpolation of spatial data can be of tremendous value in various applications, such as forecasting weather from only a few measurements of meteorological or remote sensing data. Existing methods for spatial interpolation, such as variants of kriging and spatial autoregressive models, tend to suffer from at least one of the following limitations:
openaire +2 more sources
Computing a Mechanism for a Bayesian and Partially Observable Markov Approach
The design of incentive-compatible mechanisms for a certain class of finite Bayesian partially observable Markov games is proposed using a dynamic framework.
Clempner Julio B., Poznyak Alexander S.
doaj +1 more source
A Convex Programming Approach for Discrete-Time Markov Decision Processes under the Expected Total Reward Criterion [PDF]
In this work, we study discrete-time Markov decision processes (MDPs) under constraints with Borel state and action spaces and where all the performance functions have the same form of the expected total reward (ETR) criterion over the infinite time horizon. One of our objective is to propose a convex programming formulation for this type of MDPs.
Dufour, François, Genadot, Alexandre
openaire +4 more sources
Learn Quasi-Stationary Distributions of Finite State Markov Chain
We propose a reinforcement learning (RL) approach to compute the expression of quasi-stationary distribution. Based on the fixed-point formulation of quasi-stationary distribution, we minimize the KL-divergence of two Markovian path distributions induced
Zhiqiang Cai, Ling Lin, Xiang Zhou
doaj +1 more source
A Linear Programming Approach to Markov Reward Error Bounds for Queueing Networks [PDF]
In this paper, we present a numerical framework for constructing bounds on stationary performance measures of random walks in the positive orthant using the Markov reward approach. These bounds are established in terms of stationary performance measures of a perturbed random walk whose stationary distribution is known explicitly.
Xinwei Bai, Jasper Goseling
openaire +2 more sources
A Self-Driving Decision Making With Reachable Path Analysis and Interaction-Aware Speed Profiling
This paper proposes a behavior planning algorithm for self-driving vehicles to handle lane keeping, speed control considering inter-vehicle space, and collision avoidance under uncertainty. The behavior planning approach is structured as a hierarchically
Yuho Song, Sangwon Han, Kunsoo Huh
doaj +1 more source
Strategy Selection and Outcome Evaluation of Three-Way Decisions Based on Reinforcement Learning [PDF]
The trisecting-acting-outcome (TAO) model of three-way decision (3WD) consists of three steps: trisect a whole, design action strategies, and outcome analysis and measurement.
LIU Xiaoxue, JIANG Chunmao
doaj +1 more source
Multiphase until formulas over Markov reward models: An algebraic approach
zbMATH Open Web Interface contents unavailable due to conflicting licenses.
Ming Xu 0010 +4 more
openaire +3 more sources
Dopamine, reward learning, and active inference
Temporal difference learning models propose phasic dopamine signalling encodes reward prediction errors that drive learning. This is supported by studies where optogenetic stimulation of dopamine neurons can stand in lieu of actual reward.
Thomas eFitzgerald +3 more
doaj +1 more source
A Bayesian Account of Generalist and Specialist Formation Under the Active Inference Framework
This paper offers a formal account of policy learning, or habitual behavioral optimization, under the framework of Active Inference. In this setting, habit formation becomes an autodidactic, experience-dependent process, based upon what the agent sees ...
Anthony G. Chen +4 more
doaj +1 more source

