Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives

Hüyük, A.; Tekin, Cem

Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives

buir.contributor.author	Tekin, Cem
buir.contributor.orcid	Tekin, Cem\|0000-0003-4361-4021
dc.citation.epage	1266	en_US
dc.citation.spage	1233	en_US
dc.citation.volumeNumber	110	en_US
dc.contributor.author	Hüyük, A.
dc.contributor.author	Tekin, Cem
dc.date.accessioned	2022-02-09T10:46:04Z
dc.date.available	2022-02-09T10:46:04Z
dc.date.issued	2021-06
dc.department	Department of Electrical and Electronics Engineering	en_US
dc.description.abstract	We consider multi-objective multi-armed bandit with (i) lexicographically ordered and (ii) satisficing objectives. In the first problem, the goal is to select arms that are lexicographic optimal as much as possible without knowing the arm reward distributions beforehand. We capture this goal by defining a multi-dimensional form of regret that measures the loss due to not selecting lexicographic optimal arms, and then, propose an algorithm that achieves O~(T2/3) gap-free regret and prove a regret lower bound of Ω(T2/3). We also consider two additional settings where the learner has prior information on the expected arm rewards. In the first setting, the learner only knows for each objective the lexicographic optimal expected reward. In the second setting, it only knows for each objective a near-lexicographic optimal expected reward. For both settings, we prove that the learner achieves expected regret uniformly bounded in time. Then, we show that the algorithm we propose for the second setting of lexicographically ordered objectives with prior information also attains bounded regret for satisficing objectives. Finally, we experimentally evaluate the proposed algorithms in a variety of multi-objective learning problems.	en_US
dc.identifier.doi	10.1007/s10994-021-05956-1	en_US
dc.identifier.eissn	1573-0565
dc.identifier.issn	0885-6125
dc.identifier.uri	http://hdl.handle.net/11693/77165
dc.language.iso	English	en_US
dc.publisher	Springer	en_US
dc.relation.isversionof	https://doi.org/10.1007/s10994-021-05956-1	en_US
dc.source.title	Machine Learning	en_US
dc.subject	Multi-armed bandit	en_US
dc.subject	Multi-objective learning	en_US
dc.subject	Lexicographic optimality	en_US
dc.subject	Satisficing	en_US
dc.title	Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives	en_US
dc.type	Article	en_US

Files

Original bundle

Now showing 1 - 1 of 1

Name:: Multi-objective_multi-armed_bandit_with_lexicographically_ordered_and_satisficing_objectives.pdf
Size:: 3.82 MB
Format:: Adobe Portable Document Format
Description:

Download

License bundle

Now showing 1 - 1 of 1

Name:: license.txt
Size:: 1.69 KB
Format:: Item-specific license agreed upon to submission
Description:

Download

Collections

Scholarly Publications - Electrical and Electronics Engineering