In this paper we propose a new methodology for solving a discrete time\nstochastic Markovian control problem under model uncertainty. By utilizing the\nDirichlet process, we model the unknown distribution of the underlying\nstochastic process as a random probability measure and achieve online learning\nin a Bayesian manner. Our approach integrates optimizing and dynamic learning.\nWhen dealing with model uncertainty, the nonparametric framework allows us to\navoid model misspecification that usually occurs in other classical control\nmethods. Then, we develop a numerical algorithm to handle the infinitely\ndimensional state space in this setup and utilizes Gaussian process surrogates\nto obtain a functional representation of the value function in the Bellman\nrecursion. We also build separate surrogates for optimal control to eliminate\nrepeated optimizations on out-of-sample paths and bring computational\nspeed-ups. Finally, we demonstrate the financial advantages of the\nnonparametric Bayesian framework compared to parametric approaches such as\nstrong robust and time consistent adaptive.\n