Model-free Reinforcement Learning for Branching Markov Decision Processes

We study reinforcement learning for the optimal control of Branching Markov\nDecision Processes (BMDPs), a natural extension of (multitype) Branching Markov\nChains (BMCs). The state of a (discrete-time) BMCs is a collection of entities\nof various types that, while spawning other entities, generate a payoff. In\ncomparison with BMCs, where the evolution of a each entity of the same type\nfollows the same probabilistic pattern, BMDPs allow an external controller to\npick from a range of options. This permits us to study the best/worst behaviour\nof the system. We generalise model-free reinforcement learning techniques to\ncompute an optimal control strategy of an unknown BMDP in the limit. We present\nresults of an implementation that demonstrate the practicality of the approach.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC