We introduce a simple new method for visual imitation learning, which allows\na novel robot manipulation task to be learned from a single human\ndemonstration, without requiring any prior knowledge of the object being\ninteracted with. Our method models imitation learning as a state estimation\nproblem, with the state defined as the end-effector's pose at the point where\nobject interaction begins, as observed from the demonstration. By then\nmodelling a manipulation task as a coarse, approach trajectory followed by a\nfine, interaction trajectory, this state estimator can be trained in a\nself-supervised manner, by automatically moving the end-effector's camera\naround the object. At test time, the end-effector moves to the estimated state\nthrough a linear path, at which point the original demonstration's end-effector\nvelocities are simply replayed. This enables convenient acquisition of a\ncomplex interaction trajectory, without actually needing to explicitly learn a\npolicy. Real-world experiments on 8 everyday tasks show that our method can\nlearn a diverse range of skills from a single human demonstration, whilst also\nyielding a stable and interpretable controller.\n