Results 31 to 40 of about 6,497 (242)
Approximate receding horizon approach for Markov decision processes: average reward case
The authors consider an approximation scheme for solving Markov decision processes (MDPs) with countable state space, finite action space, and bounded rewards that uses an approximate solution of a fixed finite-horizon sub-MDP of a given infinite-horizon MDP to create a stationary policy, which they call ''approximate receding horizon control''.
Chang, Hyeong Soo, Marcus, Steven I.
openaire +2 more sources
Object Affordance Driven Inverse Reinforcement Learning Through Conceptual Abstraction and Advice
Within human Intent Recognition (IR), a popular approach to learning from demonstration is Inverse Reinforcement Learning (IRL). IRL extracts an unknown reward function from samples of observed behaviour. Traditional IRL systems require large datasets to
Bhattacharyya Rupam +1 more
doaj +1 more source
A State Aggregation Approach To Singularly Perturbed Markov Reward Processes
In this paper, we propose a single sample path based algorithm with state aggregation to optimize the average rewards of singularly perturbed Markov reward processes (SPMRPs) with a large scale state spaces. It is assumed that such a reward process depend on a set of parameters.
Zhang, Dali, Baoqun Yin, Hongsheng Xi
openaire +2 more sources
Learning Highly Dynamic Skills Transition for Quadruped Jumping Through Constrained Space
A quadruped robot masters dynamic jumps through constrained spaces with animal‐inspired moves and intelligent vision control. This hierarchical learning approach combines imitation of biological agility with real‐time trajectory planning. Although legged animals are capable of performing explosive motions while traversing confined spaces, replicating ...
Zeren Luo +6 more
wiley +1 more source
This work presents a robotic control method for human–robot collaborative assembly based on a biomechanics‐constrained digital human model. Reinforcement learning is used to generate physiologically plausible human motion trajectories, which are integrated into a virtual environment for robot control learning.
Bitao Yao +4 more
wiley +1 more source
Optimizing Subway Train Operation With Hierarchical Adaptive Control Approach
The proportional integral derivative (PID) method is widely used in industrial control applications. However, when applied to complex and dynamic train operation control systems, real-time parameter adjustment becomes a formidable challenge.
Gaoyun Cheng +6 more
doaj +1 more source
Single‐cell longitudinal profiling reveals that androgen‐deprivation therapy induces a DPT+ fibroblast‐complement axis that suppresses macrophage inflammation and drives CD8+ T cell exhaustion in prostate cancer. Concurrently, resistant epithelial subpopulations persist and engage TSPAN1‐ and NRXN1‐mediated programs promoting CRPC and neuroendocrine ...
Yang Chen +19 more
wiley +1 more source
Sustainable Materials Design With Multi‐Modal Artificial Intelligence
Critical mineral scarcity, high embodied carbon, and persistent pollution from materials processing intensify the need for sustainable materials design. This review frames the problem as multi‐objective optimization under heterogeneous, high‐dimensional evidence and highlights multi‐modal AI as an enabling pathway.
Tianyi Xu +8 more
wiley +1 more source
Modified Index Policies for Multi-Armed Bandits with Network-like Markovian Dependencies
Sequential decision-making in dynamic and interconnected environments is a cornerstone of numerous applications, ranging from communication networks and finance to distributed blockchain systems and IoT frameworks. The multi-armed bandit (MAB) problem is
Abdalaziz Sawwan, Jie Wu
doaj +1 more source
This paper is concerned with the asymptotic optimality of quantized stationary policies for continuous-time Markov decision processes (CTMDPs) in Polish spaces with state-dependent discount factors, where the transition rates and reward rates are allowed
Xiao Wu, Yanqiu Tang
doaj +1 more source

