Results 21 to 30 of about 6,497 (242)

Indirect Dynamic Negotiation in the Nash Demand Game

open access: yesIEEE Access, 2022
The paper addresses a problem of sequential bilateral bargaining with incomplete information. We proposed a decision model that helps agents to successfully bargain by performing indirect negotiation and learning the opponent’s model ...
Tatiana V. Guy   +2 more
doaj   +1 more source

A Markov chain Monte Carlo algorithm for Bayesian policy search

open access: yesSystems Science & Control Engineering, 2018
Policy search algorithms have facilitated application of Reinforcement Learning (RL) to dynamic systems, such as control of robots. Many policy search algorithms are based on the policy gradient, and thus may suffer from slow convergence or local optima ...
Vahid Tavakol Aghaei   +2 more
doaj   +1 more source

Inverse reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units

open access: yesBMC Medical Informatics and Decision Making, 2019
Background Reinforcement learning (RL) provides a promising technique to solve complex sequential decision making problems in health care domains. To ensure such applications, an explicit reward function encoding domain knowledge should be specified ...
Chao Yu, Jiming Liu, Hongyi Zhao
doaj   +1 more source

A Reinforcement Learning-Based Congestion Control Approach for V2V Communication in VANET

open access: yesApplied Sciences, 2023
Vehicular ad hoc networks (VANETs) are crucial components of intelligent transportation systems (ITS) aimed at enhancing road safety and providing additional services to vehicles and their users.
Xiaofeng Liu   +2 more
doaj   +1 more source

Bounds on the bias terms for the Markov reward approach

open access: yes, 2019
An important step in the Markov reward approach to error bounds on stationary performance measures of Markov chains is to bound the bias terms. Affine functions have been successfully used for these bounds for various models, but there are also models for which it has not been possible to establish such bounds.
Bai, Xinwei, Goseling, Jasper
openaire   +2 more sources

Differentially Private Reward Functions in Markov Decision Processes: Policy Synthesis and Tradeoffs

open access: yesIEEE Open Journal of Control Systems
Policy synthesis in Markov decision processes uses a known reward function to compute a policy that maximizes it. However, onlookers may infer reward functions by observing agents, which can reveal sensitive information.
Alexander Benvenuti   +6 more
doaj   +1 more source

Research on Wargame Decision-Making Method Based on Multi-Agent Deep Deterministic Policy Gradient

open access: yesApplied Sciences, 2023
Wargames are essential simulators for various war scenarios. However, the increasing pace of warfare has rendered traditional wargame decision-making methods inadequate.
Sheng Yu, Wei Zhu, Yong Wang
doaj   +1 more source

Implicit Value Updating Explains Transitive Inference Performance: The Betasort Model. [PDF]

open access: yesPLoS Computational Biology, 2015
Transitive inference (the ability to infer that B > D given that B > C and C > D) is a widespread characteristic of serial learning, observed in dozens of species.
Greg Jensen   +4 more
doaj   +1 more source

Peer-to-peer trading in smart grid with demand response and grid outage using deep reinforcement learning

open access: yesAin Shams Engineering Journal, 2023
With the price of green energy now more reasonable, users can now produce enough electricity to meet their needs and make a profit by selling the surplus on the underground P2P energy market.
Mohammed Alsolami   +3 more
doaj   +1 more source

Dynamic Mechanism Design for Repeated Markov Games with Hidden Actions: Computational Approach

open access: yesMathematical and Computational Applications
This paper introduces a dynamic mechanism design tailored for uncertain environments where incentive schemes are challenged by the inability to observe players’ actions, known as moral hazard.
Julio B. Clempner
doaj   +1 more source

Home - About - Disclaimer - Privacy