This paper presents MAST, a new model for Multimodal Abstractive Text\nSummarization that utilizes information from all three modalities -- text,\naudio and video -- in a multimodal video. Prior work on multimodal abstractive\ntext summarization only utilized information from the text and video\nmodalities. We examine the usefulness and challenges of deriving information\nfrom the audio modality and present a sequence-to-sequence trimodal\nhierarchical attention-based model that overcomes these challenges by letting\nthe model pay more attention to the text modality. MAST outperforms the current\nstate of the art model (video-text) by 2.51 points in terms of Content F1 score\nand 1.00 points in terms of Rouge-L score on the How2 dataset for multimodal\nlanguage understanding.\n