Results 31 to 40 of about 18,568,643 (214)
Learning the distribution with largest mean: two bandit frameworks*
Over the past few years, the multi-armed bandit model has become increasingly popular in the machine learning community, partly because of applications including online content optimization. This paper reviews two different sequential learning tasks that
Kaufmann Emilie, Garivier Aurélien
doaj +1 more source
Ellsberg paradox in decision theory posits that people will inevitably choose a known probability of winning over an unknown probability of winning even if the known probability is low [1].
Song-Ju Kim, Taiki Takahashi
doaj +1 more source
Fast Two-Stage Computation of an Index Policy for Multi-Armed Bandits with Setup Delays
We consider the multi-armed bandit problem with penalties for switching that include setup delays and costs, extending the former results of the author for the special case with no switching delays.
José Niño-Mora
doaj +1 more source
One of two independent Bernoulli processes (arms) with unknown expectations $\rho$ and $\lambda$ is selected and observed at each of $n$ stages. The selection problem is sequential in that the process which is selected at a particular stage is a function of the results of previous selections as well as of prior information about $\rho$ and $\lambda ...
openaire +2 more sources
Strategic Experimentation with Private Payoffs [PDF]
We consider two players facing identical discrete-time bandit problems with a safe and a risky arm. In any period, the risky arm yields either a success or a failure, and the first success reveals the risky arm to dominate the safe one.
Heidhues, Paul +2 more
core +1 more source
Multi-armed bandit-based adaptive control of advertising in social networks
In a rapidly changing marketing environment, to ensure the success of an advertising campaign, it is necessary to minimize the time from the idea of advertising to the implementation of communication.
О. Shtovba
doaj +1 more source
Randomization in the Two-Armed Bandit Problem
Keywords: randomization ; existence of optimal solutions ; continuous-time two-armed bandit Reference PROB-ARTICLE-1990-001doi:10.1214/aop/1176990946 Record created on 2008-12-01, modified on 2017-05 ...
openaire +2 more sources
Highly Configurable Evolutionary Multi‐Objective Optimisation and Application in UAV Path Planning
ABSTRACT Multi‐objective optimisation is crucial in engineering and real‐world decision‐making. Although traditional decomposition‐based approaches such as MOEA/D yield reasonable results, they depend heavily on expert knowledge. Existing MetaBBO approaches partially alleviate this reliance, however, they offer limited configurability and lack real ...
Kaixu Chen +4 more
wiley +1 more source
The Two-Armed Bandit with Delayed Responses
A general model for a two-armed bandit with delayed responses is introduced and solved with dynamic programming. One arm has geometric lifetime with parameter \(\theta\), which has prior distribution \(\mu\). The other arm has known lifetime with mean \(\kappa\).
openaire +2 more sources
ABSTRACT The British army in India took great care to provide European troops with facilities for sexual relations while anxiously managing venereal disease. Examining archival evidence, political debates and medical discourse from the nineteenth century, this article examines the colonial military enterprise of regulated prostitution in colonial ...
Sameera Chauhan
wiley +1 more source

