Video anomaly detection has proved to be a challenging task owing to its\nunsupervised training procedure and high spatio-temporal complexity existing in\nreal-world scenarios. In the absence of anomalous training samples,\nstate-of-the-art methods try to extract features that fully grasp normal\nbehaviors in both space and time domains using different approaches such as\nautoencoders, or generative adversarial networks. However, these approaches\ncompletely ignore or, by using the ability of deep networks in the hierarchical\nmodeling, poorly model the spatio-temporal interactions that exist between\nobjects. To address this issue, we propose a novel yet efficient method named\nAno-Graph for learning and modeling the interaction of normal objects. Towards\nthis end, a Spatio-Temporal Graph (STG) is made by considering each node as an\nobject's feature extracted from a real-time off-the-shelf object detector, and\nedges are made based on their interactions. After that, a self-supervised\nlearning method is employed on the STG in such a way that encapsulates\ninteractions in a semantic space. Our method is data-efficient, significantly\nmore robust against common real-world variations such as illumination, and\npasses SOTA by a large margin on the challenging datasets ADOC and Street Scene\nwhile stays competitive on Avenue, ShanghaiTech, and UCSD.\n