Exploiting Side Information for Improved Online Learning Algorithms in Wireless Networks
Manjesh K. Hanawal, Sumit J. Darak · IEEE Transactions on Wireless Communications · 2024
In wireless networks, the transmitter adapts its parameters based on the receiver’s feedback to achieve a high throughput. The throughput also depends on factors like interference level and channel gain, which can be measured at the transmitter. They provide useful information about the instantaneous throughput as well. For example, higher interference implies a lower throughput. This work treats any measurable quality with a non-zero correlation with the throughput as side information (SI). We also study how it can be exploited to quickly learn the channel that offers higher throughput (reward). When the mean value of the SI is known, using control variate theory, we develop online learning algorithms that require fewer samples to learn and can improve the learning rate compared to cases where SI is ignored. Specifically, we incorporated SI in the Upper Confidence Bound (UCB) algorithm and proposed the UCBwSI algorithm. We quantify the gain achieved in terms of the regret and show that the improvement in regret over state-of-the-art UCB is proportional to the correlation between the reward and SI. Simulations demonstrate a 5-10% improvement in the bit-error rate. Even when the mean of the SI is unknown, we demonstrate the superiority of the UCBwSI over UCB.