The biobjective multiarmed bandit: learning approximate lexicographic optimal allocations

Tekin, Cem

The biobjective multiarmed bandit: learning approximate lexicographic optimal allocations

buir.contributor.author	Tekin, Cem
dc.citation.epage	1080	en_US
dc.citation.issueNumber	2	en_US
dc.citation.spage	1065	en_US
dc.citation.volumeNumber	27	en_US
dc.contributor.author	Tekin, Cem	en_US
dc.date.accessioned	2020-02-24T07:45:50Z
dc.date.available	2020-02-24T07:45:50Z
dc.date.issued	2019	en_US
dc.department	Department of Electrical and Electronics Engineering	en_US
dc.description.abstract	We consider a biobjective sequential decision-making problem where an allocation (arm) is called ϵ lexicographic optimal if its expected reward in the first objective is at most ϵ smaller than the highest expected reward, and its expected reward in the second objective is at least the expected reward of a lexicographic optimal arm. The goal of the learner is to select arms that are ϵ lexicographic optimal as much as possible without knowing the arm reward distributions beforehand. For this problem, we first show that the learner’s goal is equivalent to minimizing the ϵ lexicographic regret, and then, propose a learning algorithm whose ϵ lexicographic gap-dependent regret is bounded and gap-independent regret is sublinear in the number of rounds with high probability. Then, we apply the proposed model and algorithm for dynamic rate and channel selection in a cognitive radio network with imperfect channel sensing. Our results show that the proposed algorithm is able to learn the approximate lexicographic optimal rate–channel pair that simultaneously minimizes the primary user interference and maximizes the secondary user throughput.	en_US
dc.identifier.doi	10.3906/elk-1806-221	en_US
dc.identifier.issn	1300-0632
dc.identifier.uri	http://hdl.handle.net/11693/53477
dc.language.iso	English	en_US
dc.publisher	TÜBİTAK	en_US
dc.relation.isversionof	https://dx.doi.org/10.3906/elk-1806-221	en_US
dc.source.title	Turkish Journal of Electrical Engineering and Computer Sciences	en_US
dc.subject	Multiarmed bandit	en_US
dc.subject	Biobjective learning	en_US
dc.subject	Lexicographic optimality	en_US
dc.subject	Dynamic rate and channel selection	en_US
dc.subject	Cognitive radio networks	en_US
dc.title	The biobjective multiarmed bandit: learning approximate lexicographic optimal allocations	en_US
dc.type	Article	en_US

Files

Original bundle

Now showing 1 - 1 of 1

Name:: The_biobjective_multiarmed_bandit_Learning_approximate_lexicographic_optimal_allocations.pdf
Size:: 313.19 KB
Format:: Adobe Portable Document Format
Description:

Download

License bundle

Now showing 1 - 1 of 1

Name:: license.txt
Size:: 1.71 KB
Format:: Item-specific license agreed upon to submission
Description:

Download

Collections

Scholarly Publications - Electrical and Electronics Engineering