We propose detection of deepfake videos based on the dissimilarity between\nthe audio and visual modalities, termed as the Modality Dissonance Score (MDS).\nWe hypothesize that manipulation of either modality will lead to dis-harmony\nbetween the two modalities, eg, loss of lip-sync, unnatural facial and lip\nmovements, etc. MDS is computed as an aggregate of dissimilarity scores between\naudio and visual segments in a video. Discriminative features are learnt for\nthe audio and visual channels in a chunk-wise manner, employing the\ncross-entropy loss for individual modalities, and a contrastive loss that\nmodels inter-modality similarity. Extensive experiments on the DFDC and\nDeepFake-TIMIT Datasets show that our approach outperforms the state-of-the-art\nby up to 7%. We also demonstrate temporal forgery localization, and show how\nour technique identifies the manipulated video segments.\n