Results 31 to 40 of about 311 (117)

Modeling sensory-motor decisions in natural behavior. [PDF]

open access: yesPLoS Computational Biology, 2018
Although a standard reinforcement learning model can capture many aspects of reward-seeking behaviors, it may not be practical for modeling human natural behaviors because of the richness of dynamic environments and limitations in cognitive resources. We
Ruohan Zhang   +6 more
doaj   +1 more source

Generative Adversarial Inverse Reinforcement Learning With Deep Deterministic Policy Gradient

open access: yesIEEE Access, 2023
Although the issue of sparse expert samples at the early stage of training in inverse reinforcement learning (IRL) is successfully resolved by the introduction of generative adversarial network (GAN), the inherent drawbacks of GAN result in ineffective ...
Ming Zhan, Jingjing Fan, Jianying Guo
doaj   +1 more source

Off-Dynamics Inverse Reinforcement Learning

open access: yesIEEE Access
Imitation learning is a widely-used paradigm for decision making that learns from expert demonstrations. Existing imitation algorithms often require multiple interactions between the agent and the environment from which the demonstration is obtained. The
Yachen Kang, Jinxin Liu, Donglin Wang
doaj   +1 more source

Bayesian Multitask Inverse Reinforcement Learning [PDF]

open access: yes, 2012
We generalise the problem of inverse reinforcement learning to multiple tasks, from multiple demonstrations. Each one may represent one expert trying to solve a different task, or as different experts trying to solve the same task. Our main contribution is to formalise the problem as statistical preference elicitation, via a number of structured priors,
Christos Dimitrakakis   +1 more
openaire   +2 more sources

Preference Elicitation and Inverse Reinforcement Learning [PDF]

open access: yes, 2011
We state the problem of inverse reinforcement learning in terms of preference elicitation, resulting in a principled (Bayesian) statistical formulation. This generalises previous work on Bayesian inverse reinforcement learning and allows us to obtain a posterior distribution on the agent's preferences, policy and optionally, the obtained reward ...
Constantin A. Rothkopf   +1 more
openaire   +3 more sources

Inverse Delayed Reinforcement Learning

open access: yesCoRR
Inverse Reinforcement Learning (IRL) has demonstrated effectiveness in a variety of imitation tasks. In this paper, we introduce an IRL framework designed to extract rewarding features from expert trajectories affected by delayed disturbances. Instead of relying on direct observations, our approach employs an efficient off-policy adversarial training ...
Simon Sinong Zhan   +8 more
openaire   +2 more sources

Integrating inverse reinforcement learning into data-driven mechanistic computational models: a novel paradigm to decode cancer cell heterogeneity

open access: yesFrontiers in Systems Biology
Cellular heterogeneity is a ubiquitous aspect of biology and a major obstacle to successful cancer treatment. Several techniques have emerged to quantify heterogeneity in live cells along axes including cellular migration, morphology, growth, and ...
Patrick C. Kinnunen   +17 more
doaj   +1 more source

Distributional Inverse Reinforcement Learning

open access: yesCoRR
ICML 2026 ...
Feiyang Wu, Ye Zhao, Anqi Wu
openaire   +2 more sources

Gamma-Regression-Based Inverse Reinforcement Learning From Suboptimal Demonstrations

open access: yesIEEE Access
Inverse reinforcement learning (IRL) is a technique that estimates the intention of an expert who acts optimally on a specific intention, as a reward from demonstration (i.e., recorded data of the expert’s behavior). Traditional IRL algorithms are
Daiko Kishikawa, Sachiyo Arai
doaj   +1 more source

Variational Reward Estimator Bottleneck: Towards Robust Reward Estimator for Multidomain Task-Oriented Dialogue

open access: yesApplied Sciences, 2021
Despite its significant effectiveness in adversarial training approaches to multidomain task-oriented dialogue systems, adversarial inverse reinforcement learning of the dialogue policy frequently fails to balance the performance of the reward estimator ...
Jeiyoon Park   +4 more
doaj   +1 more source

Home - About - Disclaimer - Privacy