Online Restless Multi-Armed Bandits with Long-Term Fairness Constraints
Restless multi-armed bandits (RMAB) have been widely used to model sequential decision making problems with constraints. The decision maker (DM) aims to maximize the expected total reward over an infinite horizon under an “instantaneous activation constraint” that at most B arms can be activated at any decision epoch, where the state of each arm ...
Shufan Wang, Guojun Xiong, Jian Li
openaire +2 more sources
Multi-Armed Bandits in Brain-Computer Interfaces. [PDF]
Heskebeck F +2 more
europepmc +1 more source
Signal detection models as contextual bandits. [PDF]
Sherratt TN, O'Neill E.
europepmc +1 more source
Optimal Control of Fluid Restless Multi-armed Bandits: A Machine Learning Approach
We present a novel machine learning framework for the optimal control of fluid restless multi-armed bandit problems (FRMABPs) with state equations that are either affine or quadratic in the state variables. By establishing fundamental properties of FRMABPs, we develop an efficient numerical algorithm that generates a comprehensive training set by ...
Dimitris Bertsimas +2 more
openaire +2 more sources
Reliability of Decision-Making and Reinforcement Learning Computational Parameters. [PDF]
Mkrtchian A, Valton V, Roiser JP.
europepmc +1 more source
Minimizing Cost Rather Than Maximizing Reward in Restless Multi-Armed Bandits
Restless Multi-Armed Bandits (RMABs) offer a powerful framework for solving resource constrained maximization problems. However, the formulation can be inappropriate for settings where the limiting constraint is a reward threshold rather than a budget.
R. Teal Witter, Lisa Hellerstein
openaire +2 more sources
Contextual Restless Multi-Armed Bandits with Application to Demand Response Decision-Making
This paper introduces a novel multi-armed bandits framework, termed Contextual Restless Bandits (CRB), for complex online decision-making. This CRB framework incorporates the core features of contextual bandits and restless bandits, so that it can model both the internal state transitions of each arm and the influence of external global environmental ...
Xin Chen, I-Hong Hou
openaire +2 more sources
Attenuated Directed Exploration during Reinforcement Learning in Gambling Disorder. [PDF]
Wiehler A, Chakroun K, Peters J.
europepmc +1 more source
Dynamic Content Caching with Waiting Costs via Restless Multi-Armed Bandits
We consider a system with a local cache connected to a backend server and an end user population. A set of contents are stored at the the server where they continuously get updated. The local cache keeps copies, potentially stale, of a subset of the contents. The users make content requests to the local cache which either can serve the local version if
Ankita Koley, Chandramani Singh
openaire +2 more sources
Restless Multi-Process Multi-Armed Bandits with Applications to Self-Driving Microscopies
High-content screening microscopy generates large amounts of live-cell imaging data, yet its potential remains constrained by the inability to determine when and where to image most effectively. Optimally balancing acquisition time, computational capacity, and photobleaching budgets across thousands of dynamically evolving regions of interest remains ...
Jaume Anguera Peris +4 more
openaire +2 more sources

