In order for reinforcement learning techniques to be useful in real-world\ndecision making processes, they must be able to produce robust performance from\nlimited data. Deep policy optimization methods have achieved impressive results\non complex tasks, but their real-world adoption remains limited because they\noften require significant amounts of data to succeed. When combined with small\nsample sizes, these methods can result in unstable learning due to their\nreliance on high-dimensional sample-based estimates. In this work, we develop\ntechniques to control the uncertainty introduced by these estimates. We\nleverage these techniques to propose a deep policy optimization approach\ndesigned to produce stable performance even when data is scarce. The resulting\nalgorithm, Uncertainty-Aware Trust Region Policy Optimization, generates robust\npolicy updates that adapt to the level of uncertainty present throughout the\nlearning process.\n