Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives

Hüyük, A.; Tekin, Cem

Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives

Files

Multi-objective_multi-armed_bandit_with_lexicographically_ordered_and_satisficing_objectives.pdf (3.82 MB)

Date

2021-06

Authors

Hüyük, A.

Tekin, Cem

BUIR Usage Stats

2
views

9
downloads

Citation Stats

Attention Stats

Abstract

We consider multi-objective multi-armed bandit with (i) lexicographically ordered and (ii) satisficing objectives. In the first problem, the goal is to select arms that are lexicographic optimal as much as possible without knowing the arm reward distributions beforehand. We capture this goal by defining a multi-dimensional form of regret that measures the loss due to not selecting lexicographic optimal arms, and then, propose an algorithm that achieves O~(T2/3) gap-free regret and prove a regret lower bound of Ω(T2/3). We also consider two additional settings where the learner has prior information on the expected arm rewards. In the first setting, the learner only knows for each objective the lexicographic optimal expected reward. In the second setting, it only knows for each objective a near-lexicographic optimal expected reward. For both settings, we prove that the learner achieves expected regret uniformly bounded in time. Then, we show that the algorithm we propose for the second setting of lexicographically ordered objectives with prior information also attains bounded regret for satisficing objectives. Finally, we experimentally evaluate the proposed algorithms in a variety of multi-objective learning problems.

Source Title

Machine Learning

Publisher

Springer

Keywords

Multi-armed bandit, Multi-objective learning, Lexicographic optimality, Satisficing

Permalink

http://hdl.handle.net/11693/77165

Published Version (Please cite this version)

https://doi.org/10.1007/s10994-021-05956-1

Collections

Scholarly Publications - Electrical and Electronics Engineering

Language

English

Type

Article

Full item page

Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives

Files

Date

Authors

Editor(s)

Advisor

Supervisor

Co-Advisor

Co-Supervisor

Instructor

BUIR Usage Stats

Citation Stats

Attention Stats

Series

Abstract

Source Title

Publisher

Course

Other identifiers

Book Title

Keywords

Degree Discipline

Degree Level

Degree Name

Citation

Permalink

Published Version (Please cite this version)

Collections

Language

Type

Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives

Files

Date

Authors

Editor(s)

Advisor

Supervisor

Co-Advisor

Co-Supervisor

Instructor

BUIR Usage Stats

Citation Stats

Attention Stats

Share

Series

Abstract

Source Title

Publisher

Course

Other identifiers

Book Title

Keywords

Degree Discipline

Degree Level

Degree Name

Citation

Permalink

Published Version (Please cite this version)

Collections

Language

Type