Results 11 to 20 of about 6,497 (242)

A Markov Reward Process-Based Approach to Spatial Interpolation

open access: yesCoRR, 2021
The interpolation of spatial data can be of tremendous value in various applications, such as forecasting weather from only a few measurements of meteorological or remote sensing data. Existing methods for spatial interpolation, such as variants of kriging and spatial autoregressive models, tend to suffer from at least one of the following limitations:
openaire   +2 more sources

Computing a Mechanism for a Bayesian and Partially Observable Markov Approach

open access: yesInternational Journal of Applied Mathematics and Computer Science, 2023
The design of incentive-compatible mechanisms for a certain class of finite Bayesian partially observable Markov games is proposed using a dynamic framework.
Clempner Julio B., Poznyak Alexander S.
doaj   +1 more source

A Convex Programming Approach for Discrete-Time Markov Decision Processes under the Expected Total Reward Criterion [PDF]

open access: yesSIAM Journal on Control and Optimization, 2020
In this work, we study discrete-time Markov decision processes (MDPs) under constraints with Borel state and action spaces and where all the performance functions have the same form of the expected total reward (ETR) criterion over the infinite time horizon. One of our objective is to propose a convex programming formulation for this type of MDPs.
Dufour, François, Genadot, Alexandre
openaire   +4 more sources

Learn Quasi-Stationary Distributions of Finite State Markov Chain

open access: yesEntropy, 2022
We propose a reinforcement learning (RL) approach to compute the expression of quasi-stationary distribution. Based on the fixed-point formulation of quasi-stationary distribution, we minimize the KL-divergence of two Markovian path distributions induced
Zhiqiang Cai, Ling Lin, Xiang Zhou
doaj   +1 more source

A Linear Programming Approach to Markov Reward Error Bounds for Queueing Networks [PDF]

open access: yes, 2019
In this paper, we present a numerical framework for constructing bounds on stationary performance measures of random walks in the positive orthant using the Markov reward approach. These bounds are established in terms of stationary performance measures of a perturbed random walk whose stationary distribution is known explicitly.
Xinwei Bai, Jasper Goseling
openaire   +2 more sources

A Self-Driving Decision Making With Reachable Path Analysis and Interaction-Aware Speed Profiling

open access: yesIEEE Access, 2023
This paper proposes a behavior planning algorithm for self-driving vehicles to handle lane keeping, speed control considering inter-vehicle space, and collision avoidance under uncertainty. The behavior planning approach is structured as a hierarchically
Yuho Song, Sangwon Han, Kunsoo Huh
doaj   +1 more source

Strategy Selection and Outcome Evaluation of Three-Way Decisions Based on Reinforcement Learning [PDF]

open access: yesJisuanji kexue yu tansuo
The trisecting-acting-outcome (TAO) model of three-way decision (3WD) consists of three steps: trisect a whole, design action strategies, and outcome analysis and measurement.
LIU Xiaoxue, JIANG Chunmao
doaj   +1 more source

Multiphase until formulas over Markov reward models: An algebraic approach

open access: yesTheoretical Computer Science, 2016
zbMATH Open Web Interface contents unavailable due to conflicting licenses.
Ming Xu 0010   +4 more
openaire   +3 more sources

Dopamine, reward learning, and active inference

open access: yesFrontiers in Computational Neuroscience, 2015
Temporal difference learning models propose phasic dopamine signalling encodes reward prediction errors that drive learning. This is supported by studies where optogenetic stimulation of dopamine neurons can stand in lieu of actual reward.
Thomas eFitzgerald   +3 more
doaj   +1 more source

A Bayesian Account of Generalist and Specialist Formation Under the Active Inference Framework

open access: yesFrontiers in Artificial Intelligence, 2020
This paper offers a formal account of policy learning, or habitual behavioral optimization, under the framework of Active Inference. In this setting, habit formation becomes an autodidactic, experience-dependent process, based upon what the agent sees ...
Anthony G. Chen   +4 more
doaj   +1 more source

Home - About - Disclaimer - Privacy