Reinforcement Learning in Linear Quadratic Deep Structured Teams: Global Convergence of Policy Gradient Methods

In this paper, we study the global convergence of model-based and model-free\npolicy gradient descent and natural policy gradient descent algorithms for\nlinear quadratic deep structured teams. In such systems, agents are partitioned\ninto a few sub-populations wherein the agents in each sub-population are\ncoupled in the dynamics and cost function through a set of linear regressions\nof the states and actions of all agents. Every agent observes its local state\nand the linear regressions of states, called deep states. For a sufficiently\nsmall risk factor and/or sufficiently large population, we prove that\nmodel-based policy gradient methods globally converge to the optimal solution.\nGiven an arbitrary number of agents, we develop model-free policy gradient and\nnatural policy gradient algorithms for the special case of risk-neutral cost\nfunction. The proposed algorithms are scalable with respect to the number of\nagents due to the fact that the dimension of their policy space is independent\nof the number of agents in each sub-population. Simulations are provided to\nverify the theoretical results.\n

Paper

References (26)

Scroll for more · 14 remaining

Similar papers

© 2026 NYSGPT2525 LLC