Explainability and interpretability of AI models is an essential factor\naffecting the safety of AI. While various explainable AI (XAI) approaches aim\nat mitigating the lack of transparency in deep networks, the evidence of the\neffectiveness of these approaches in improving usability, trust, and\nunderstanding of AI systems are still missing. We evaluate multimodal\nexplanations in the setting of a Visual Question Answering (VQA) task, by\nasking users to predict the response accuracy of a VQA agent with and without\nexplanations. We use between-subjects and within-subjects experiments to probe\nexplanation effectiveness in terms of improving user prediction accuracy,\nconfidence, and reliance, among other factors. The results indicate that the\nexplanations help improve human prediction accuracy, especially in trials when\nthe VQA system's answer is inaccurate. Furthermore, we introduce active\nattention, a novel method for evaluating causal attentional effects through\nintervention by editing attention maps. User explanation ratings are strongly\ncorrelated with human prediction accuracy and suggest the efficacy of these\nexplanations in human-machine AI collaboration tasks.\n