Output feedback linear quadratic regulator design through one-shot Q-function learning

Ahmad Pishro Asl, Ahmad Akbari · International Journal of Control · 2025

In model-free control, reinforcement learning (RL) methods such as Q-learning have been employed to design discrete-time (DT) linear quadratic regulators (LQR). However, these methods suffer from sample inefficiency due to their iterative nature. Using an alternative formulation of the LQR problem as a semi-definite program (SDP) with linear matrix inequality (LMI) constraints can address this challenge. We propose a novel approach that constructs this alternative formulation and learns the optimal output feedback Q-function exclusively from input-output data. Our method requires only a single batch of data, enabling what we term one-shot learning. Furthermore, by leveraging the Bellman equation, we achieve online controller adaptation, resulting in an off-policy, model-free, sample-efficient, and non-iterative adaptive approach that eliminates the need for an initial stabilising policy. We validate the proposed method through simulations on an unstable system and the load frequency control problem in power systems, demonstrating its effectiveness and real-world applicability.

Read the paper · More papers on PaperTik