Reinforcement Q-Learning Algorithm for H∞ Tracking Control of Unknown Discrete-Time Linear Systems

This article addresses the online reinforcement <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-learning algorithms to design <inline-formula> <tex-math notation="LaTeX">$H_{\infty }$ </tex-math></inline-formula> tracking controller for unknown discrete-time linear systems. An augmented system composed of the original system and the command generator is constructed, and a discounted performance function is introduced to establish a discounted game algebraic Riccati equation (GARE). The existence conditions of a solution to the GARE are proposed and a lower bound is found for the discount factor to assure the stability of the <inline-formula> <tex-math notation="LaTeX">$H_{\infty }$ </tex-math></inline-formula> tracking control solution. The <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-function Bellman equation is then derived, based on which the reinforcement <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-learning algorithm is developed to learn the solution to <inline-formula> <tex-math notation="LaTeX">$H_{\infty }$ </tex-math></inline-formula> tracking control problem without knowing the system dynamics. Both state-data-driven and output-data-driven reinforcement <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-learning algorithms toward finding the control policies are proposed. Unlike the value function approximation (VFA)-based approach, it is proved that the <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-learning scheme brings out no bias of solution to the <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-function Bellman equation under the probing noise satisfying the persistent excitation (PE) condition, and therefore, converges to the nominal discounted GARE solution. Moreover, the proposed output-data-driven method is more powerful than the state-data-driven method as it may not be available to completely measure the full system states in practical applications. A simulation example with a single-phase voltage-source UPS inverter is used to verify the effectiveness of the proposed <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-learning algorithms.

Paper

Full text

PDF

Reinforcement Q-Learning Algorithm for H∞ Tracking Control of Unknown Discrete-Time Linear Systems

Semantic Scholar · Engineering · 2020

Abstract

This article addresses the online reinforcement <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-learning algorithms to design <inline-formula> <tex-math notation="LaTeX">$H_{\infty }$ </tex-math></inline-formula> tracking controller for unknown discrete-time linear systems. An augmented system composed of the original system and the command generator is constructed, and a discounted performance function is introduced to establish a discounted game algebraic Riccati equation (GARE). The existence conditions of a solution to the GARE are proposed and a lower bound is found for the discount factor to assure the stability of the <inline-formula> <tex-math notation="LaTeX">$H_{\infty }$ </tex-math></inline-formula> tracking control solution. The <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-function Bellman equation is then derived, based on which the reinforcement <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-learning algorithm is developed to learn the solution to <inline-formula> <tex-math notation="LaTeX">$H_{\infty }$ </tex-math></inline-formula> tracking control problem without knowing the system dynamics. Both state-data-driven and output-data-driven reinforcement <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-learning algorithms toward finding the control policies are proposed. Unlike the value function approximation (VFA)-based approach, it is proved that the <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-learning scheme brings out no bias of solution to the <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-function Bellman equation under the probing noise satisfying the persistent excitation (PE) condition, and therefore, converges to the nominal discounted GARE solution. Moreover, the proposed output-data-driven method is more powerful than the state-data-driven method as it may not be available to completely measure the full system states in practical applications. A simulation example with a single-phase voltage-source UPS inverter is used to verify the effectiveness of the proposed <inline-formula> <tex-math notation="LaTeX">$Q$ </tex-math></inline-formula>-learning algorithms.

Similar papers

© 2026 NYSGPT2525 LLC