We present EmotiCon, a learning-based algorithm for context-aware perceived\nhuman emotion recognition from videos and images. Motivated by Frege's Context\nPrinciple from psychology, our approach combines three interpretations of\ncontext for emotion recognition. Our first interpretation is based on using\nmultiple modalities(e.g. faces and gaits) for emotion recognition. For the\nsecond interpretation, we gather semantic context from the input image and use\na self-attention-based CNN to encode this information. Finally, we use depth\nmaps to model the third interpretation related to socio-dynamic interactions\nand proximity among agents. We demonstrate the efficiency of our network\nthrough experiments on EMOTIC, a benchmark dataset. We report an Average\nPrecision (AP) score of 35.48 across 26 classes, which is an improvement of 7-8\nover prior methods. We also introduce a new dataset, GroupWalk, which is a\ncollection of videos captured in multiple real-world settings of people\nwalking. We report an AP of 65.83 across 4 categories on GroupWalk, which is\nalso an improvement over prior methods.\n