We consider the problem of efficiently learning optimal control policies and\nvalue functions over large state spaces in an online setting in which estimates\nmust be available after each interaction with the world. This paper develops an\nexplicitly model-based approach extending the Dyna architecture to linear\nfunction approximation. Dynastyle planning proceeds by generating imaginary\nexperience from the world model and then applying model-free reinforcement\nlearning algorithms to the imagined state transitions. Our main results are to\nprove that linear Dyna-style planning converges to a unique solution\nindependent of the generating distribution, under natural conditions. In the\npolicy evaluation setting, we prove that the limit point is the least-squares\n(LSTD) solution. An implication of our results is that prioritized-sweeping can\nbe soundly extended to the linear approximation case, backing up to preceding\nfeatures rather than to preceding states. We introduce two versions of\nprioritized sweeping with linear Dyna and briefly illustrate their performance\nempirically on the Mountain Car and Boyan Chain problems.\n