The ability to perceive and reason about social interactions in the context\nof physical environments is core to human social intelligence and human-machine\ncooperation. However, no prior dataset or benchmark has systematically\nevaluated physically grounded perception of complex social interactions that go\nbeyond short actions, such as high-fiving, or simple group activities, such as\ngathering. In this work, we create a dataset of physically-grounded abstract\nsocial events, PHASE, that resemble a wide range of real-life social\ninteractions by including social concepts such as helping another agent. PHASE\nconsists of 2D animations of pairs of agents moving in a continuous space\ngenerated procedurally using a physics engine and a hierarchical planner.\nAgents have a limited field of view, and can interact with multiple objects, in\nan environment that has multiple landmarks and obstacles. Using PHASE, we\ndesign a social recognition task and a social prediction task. PHASE is\nvalidated with human experiments demonstrating that humans perceive rich\ninteractions in the social events, and that the simulated agents behave\nsimilarly to humans. As a baseline model, we introduce a Bayesian inverse\nplanning approach, SIMPLE (SIMulation, Planning and Local Estimation), which\noutperforms state-of-the-art feed-forward neural networks. We hope that PHASE\ncan serve as a difficult new challenge for developing new models that can\nrecognize complex social interactions.\n