The biobjective multiarmed bandit: learning approximate lexicographic optimal allocations
buir.contributor.author | Tekin, Cem | |
dc.citation.epage | 1080 | en_US |
dc.citation.issueNumber | 2 | en_US |
dc.citation.spage | 1065 | en_US |
dc.citation.volumeNumber | 27 | en_US |
dc.contributor.author | Tekin, Cem | en_US |
dc.date.accessioned | 2020-02-24T07:45:50Z | |
dc.date.available | 2020-02-24T07:45:50Z | |
dc.date.issued | 2019 | en_US |
dc.department | Department of Electrical and Electronics Engineering | en_US |
dc.description.abstract | We consider a biobjective sequential decision-making problem where an allocation (arm) is called ϵ lexicographic optimal if its expected reward in the first objective is at most ϵ smaller than the highest expected reward, and its expected reward in the second objective is at least the expected reward of a lexicographic optimal arm. The goal of the learner is to select arms that are ϵ lexicographic optimal as much as possible without knowing the arm reward distributions beforehand. For this problem, we first show that the learner’s goal is equivalent to minimizing the ϵ lexicographic regret, and then, propose a learning algorithm whose ϵ lexicographic gap-dependent regret is bounded and gap-independent regret is sublinear in the number of rounds with high probability. Then, we apply the proposed model and algorithm for dynamic rate and channel selection in a cognitive radio network with imperfect channel sensing. Our results show that the proposed algorithm is able to learn the approximate lexicographic optimal rate–channel pair that simultaneously minimizes the primary user interference and maximizes the secondary user throughput. | en_US |
dc.description.provenance | Submitted by Evrim Ergin (eergin@bilkent.edu.tr) on 2020-02-24T07:45:50Z No. of bitstreams: 1 The_biobjective_multiarmed_bandit_Learning_approximate_lexicographic_optimal_allocations.pdf: 320706 bytes, checksum: c09270d588bd6bbc69fe4f7fcb73428e (MD5) | en |
dc.description.provenance | Made available in DSpace on 2020-02-24T07:45:50Z (GMT). No. of bitstreams: 1 The_biobjective_multiarmed_bandit_Learning_approximate_lexicographic_optimal_allocations.pdf: 320706 bytes, checksum: c09270d588bd6bbc69fe4f7fcb73428e (MD5) Previous issue date: 2019-03 | en |
dc.identifier.doi | 10.3906/elk-1806-221 | en_US |
dc.identifier.issn | 1300-0632 | |
dc.identifier.uri | http://hdl.handle.net/11693/53477 | |
dc.language.iso | English | en_US |
dc.publisher | TÜBİTAK | en_US |
dc.relation.isversionof | https://dx.doi.org/10.3906/elk-1806-221 | en_US |
dc.source.title | Turkish Journal of Electrical Engineering and Computer Sciences | en_US |
dc.subject | Multiarmed bandit | en_US |
dc.subject | Biobjective learning | en_US |
dc.subject | Lexicographic optimality | en_US |
dc.subject | Dynamic rate and channel selection | en_US |
dc.subject | Cognitive radio networks | en_US |
dc.title | The biobjective multiarmed bandit: learning approximate lexicographic optimal allocations | en_US |
dc.type | Article | en_US |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- The_biobjective_multiarmed_bandit_Learning_approximate_lexicographic_optimal_allocations.pdf
- Size:
- 313.19 KB
- Format:
- Adobe Portable Document Format
- Description:
License bundle
1 - 1 of 1
No Thumbnail Available
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description: