Robotic manipulation can be formulated as inducing a sequence of spatial\ndisplacements: where the space being moved can encompass an object, part of an\nobject, or end effector. In this work, we propose the Transporter Network, a\nsimple model architecture that rearranges deep features to infer spatial\ndisplacements from visual input - which can parameterize robot actions. It\nmakes no assumptions of objectness (e.g. canonical poses, models, or\nkeypoints), it exploits spatial symmetries, and is orders of magnitude more\nsample efficient than our benchmarked alternatives in learning vision-based\nmanipulation tasks: from stacking a pyramid of blocks, to assembling kits with\nunseen objects; from manipulating deformable ropes, to pushing piles of small\nobjects with closed-loop feedback. Our method can represent complex multi-modal\npolicy distributions and generalizes to multi-step sequential tasks, as well as\n6DoF pick-and-place. Experiments on 10 simulated tasks show that it learns\nfaster and generalizes better than a variety of end-to-end baselines, including\npolicies that use ground-truth object poses. We validate our methods with\nhardware in the real world. Experiment videos and code are available at\nhttps://transporternets.github.io\n