Reinforcement Learning for Bandits with Continuous Actions and Large Context Spaces

Duckworth P.; Vallis KA.; Lacerda B.; Hawes N.

Reinforcement Learning for Bandits with Continuous Actions and Large Context Spaces

Duckworth P., Vallis KA., Lacerda B., Hawes N.

We consider the challenging scenario of contextual bandits with continuous actions and large context spaces. This is an increasingly important application area in personalised healthcare where an agent is requested to make dosing decisions based on a patient's single image scan. In this paper, we first adapt a reinforcement learning (RL) algorithm for continuous control to outperform contextual bandit algorithms specifically hand-crafted for continuous action spaces. We empirically demonstrate this on a suite of standard benchmark datasets for vector contexts. Secondly, we demonstrate that our RL agent can generalise problems with continuous actions to large context spaces, providing results that outperform previous methods on image contexts. Thirdly, we introduce a new contextual bandits test domain with multi-dimensional continuous action space and image contexts which existing tree-based methods cannot handle. We provide initial results with our RL agent.

Original publication

DOI

10.3233/FAIA230320

Type

Conference paper

Publication Date

28/09/2023

Volume

372

Pages

590 - 597

Cookies on this website

Reinforcement Learning for Bandits with Continuous Actions and Large Context Spaces

Duckworth P., Vallis KA., Lacerda B., Hawes N.

DOI

Type

Publication Date

Volume

Pages