Results 41 to 50 of about 26,957 (159)

Unifying Two Views on Multiple Mean-Payoff Objectives in Markov Decision Processes [PDF]

open access: yesLogical Methods in Computer Science, 2017
We consider Markov decision processes (MDPs) with multiple limit-average (or mean-payoff) objectives. There exist two different views: (i) the expectation semantics, where the goal is to optimize the expected mean-payoff objective, and (ii) the ...
Krishnendu Chatterjee   +2 more
doaj   +1 more source

Symblicit algorithms for optimal strategy synthesis in monotonic Markov decision processes [PDF]

open access: yesElectronic Proceedings in Theoretical Computer Science, 2014
When treating Markov decision processes (MDPs) with large state spaces, using explicit representations quickly becomes unfeasible. Lately, Wimmer et al. have proposed a so-called symblicit algorithm for the synthesis of optimal strategies in MDPs, in the
Aaron Bohy   +2 more
doaj   +1 more source

Safe Q-Learning Method Based on Constrained Markov Decision Processes

open access: yesIEEE Access, 2019
The application of reinforcement learning in industrial fields makes the safety problem of the agent a research hotspot. Traditional methods mainly alter the objective function and the exploration process of the agent to address the safety problem. Those
Yangyang Ge   +3 more
doaj   +1 more source

Life is Random, Time is Not: Markov Decision Processes with Window Objectives [PDF]

open access: yesLogical Methods in Computer Science, 2020
The window mechanism was introduced by Chatterjee et al. to strengthen classical game objectives with time bounds. It permits to synthesize system controllers that exhibit acceptable behaviors within a configurable time frame, all along their infinite ...
Thomas Brihaye   +3 more
doaj   +1 more source

On the detection of Markov decision processes

open access: yesAutomatica
We study the detection problem for a finite set of Markov decision processes (MDPs) where the MDPs have the same state and action spaces but possibly different probabilistic transition functions. Any one of these MDPs could be the model for some underlying controlled stochastic process, but it is unknown a priori which MDP is the ground truth.
Xiaoming Duan   +4 more
openaire   +2 more sources

Dynamic Watermarking for Finite Markov Decision Processes

open access: yesIEEE Open Journal of Control Systems
Dynamic watermarking is an active intrusion detection technique that can potentially detect replay attacks, spoofing attacks, and deception attacks in the feedback channel for control systems.
Jiacheng Tang   +2 more
doaj   +1 more source

Characterizing Markov Decision Processes [PDF]

open access: yes, 2002
Problem characteristics often have a significant influence on the difficulty of solving optimization problems. In this paper, we propose attributes for characterizing Markov Decision Processes (MDPs), and discuss how they affect the performance of reinforcement learning algorithms that use function approximation.
Bohdana Ratitch, Doina Precup
openaire   +1 more source

Contextual Markov Decision Processes

open access: yesCoRR, 2015
We consider a planning problem where the dynamics and rewards of the environment depend on a hidden static parameter referred to as the context. The objective is to learn a strategy that maximizes the accumulated reward across all contexts. The new model, called Contextual Markov Decision Process (CMDP), can model a customer's behavior when interacting
Assaf Hallak   +2 more
openaire   +2 more sources

One-Counter Markov Decision Processes [PDF]

open access: yesProceedings of the Twenty-First Annual ACM-SIAM Symposium on Discrete Algorithms, 2010
Updated preliminary version, submitted to ...
T. Brazdil   +4 more
openaire   +4 more sources

Creativity and Markov Decision Processes [PDF]

open access: yesCoRR
Creativity is already regularly attributed to AI systems outside specialised computational creativity (CC) communities. However, the evaluation of creativity in AI at large typically lacks grounding in creativity theory, which can promote inappropriate attributions and limit the analysis of creative behaviour.
Guckelsberger Christian   +2 more
openaire   +4 more sources

Home - About - Disclaimer - Privacy