Results 11 to 20 of about 18,568,643 (214)
A generalised gittins index for a class of multi-armed bandits with general resource requirements [PDF]
We generalise classical multiarmed bandits to allow for the distribution of a (fixed amount of a) divisible resource among the constituent bandits at each decision point.
Minty, R J, Glazebrook, K D
core +5 more sources
An Experimental Analysis of the Two-Armed Bandit Program [PDF]
We investigate, in an experimental setting, the behavior of single decision makers who at discrete time intervals over an "infinite" horizon may choose one action from a set of possible actions where this set is constant over time, i.e. a bandit problem.
Banks, Jeffrey +2 more
openaire +3 more sources
Conservative Contextual Combinatorial Cascading Bandit
Contextual combinatorial cascading bandit ( $C^{3}$ -bandit) is a powerful multi-armed bandit framework that balances the tradeoff between exploration and exploitation in the learning process.
Kun Wang
doaj +1 more source
Some Remarks on the Two-Armed Bandit [PDF]
In this paper we consider the following situation: An experimenter has to perform a total of N trial on two Bernoulli-type experiments E1 and E2 with success probabilites α and β respectively, where both α and β are unknown to him.
Fabius, J., Zwet, W. R. Van
openaire +3 more sources
Computational modeling of behavioral tasks: An illustration on a classic reinforcement learning paradigm [PDF]
There has been a growing interest among psychologists, psychiatrists and neuroscientists in applying computational modeling to behavioral data to understand animal and human behavior. Such approaches can be daunting for those without experience.
Suthaharan, Praveen +2 more
doaj +1 more source
Bandit Algorithm Driven by a Classical Random Walk and a Quantum Walk
Quantum walks (QWs) have a property that classical random walks (RWs) do not possess—the coexistence of linear spreading and localization—and this property is utilized to implement various kinds of applications.
Tomoki Yamagami +5 more
doaj +1 more source
Role of dopamine D2 receptors in optimizing choice strategy in a dynamic and uncertain environment
In order to investigate roles of dopamine receptor subtypes in reward-based learning, we examined choice behavior of dopamine D1 and D2 receptor-knockout (D1R-KO and D2R-KO, respectively) mice in an instrumental learning task with progressively ...
Shinae eKwak +5 more
doaj +1 more source
A device has two arms with unknown deterministic payoffs and the aim is to asymptotically identify the best one without spending too much time on the other. The Narendra algorithm offers a stochastic procedure to this end. We show under weak ergodic assumptions on these deterministic payoffs that the procedure eventually chooses the best arm (i.e ...
Tarrès, Pierre, Vandekerkhove, Pierre
openaire +6 more sources
On-Line Adaptation of Exploration in the One-Armed Bandit with Covariates Problem
Many sequential decision making problems require an agent to balance exploration and exploitation to maximise long-term reward. Existing policies that address this tradeoff typically have parameters that are set a priori to control the amount of ...
Sykulski, Adam M. +2 more
core +2 more sources
Theory of Acceleration of Decision-Making by Correlated Time Sequences
Photonic accelerators have been intensively studied to provide enhanced information processing capability to benefit from the unique attributes of physical processes.
Norihiro Okada +5 more
doaj +1 more source

