The lottery ticket hypothesis questions the role of overparameterization in\nsupervised deep learning. But how is the performance of winning lottery tickets\naffected by the distributional shift inherent to reinforcement learning\nproblems? In this work, we address this question by comparing sparse agents who\nhave to address the non-stationarity of the exploration-exploitation problem\nwith supervised agents trained to imitate an expert. We show that feed-forward\nnetworks trained with behavioural cloning compared to reinforcement learning\ncan be pruned to higher levels of sparsity without performance degradation.\nThis suggests that in order to solve the RL-specific distributional shift\nagents require more degrees of freedom. Using a set of carefully designed\nbaseline conditions, we find that the majority of the lottery ticket effect in\nboth learning paradigms can be attributed to the identified mask rather than\nthe weight initialization. The input layer mask selectively prunes entire input\ndimensions that turn out to be irrelevant for the task at hand. At a moderate\nlevel of sparsity the mask identified by iterative magnitude pruning yields\nminimal task-relevant representations, i.e., an interpretable inductive bias.\nFinally, we propose a simple initialization rescaling which promotes the robust\nidentification of sparse task representations in low-dimensional control tasks.\n