Recently, representation learning for text and speech has successfully\nimproved many language related tasks. However, all existing methods suffer from\ntwo limitations: (a) they only learn from one input modality, while a unified\nrepresentation for both speech and text is needed by tasks such as end-to-end\nspeech translation, and as a result,(b) they can not exploit various\nlarge-scale text and speech data and their performance is limited by the\nscarcity of parallel speech translation data.To address these problems, we\npropose a Fused Acoustic and Text Masked Language Model (FAT-MLM) which jointly\nlearns a unified representation for both acoustic and text input from various\ntypes of corpora including parallel data for speech recognition and machine\ntranslation, and even pure speech and text data. Within this cross-modal\nrepresentation learning framework, we further present an end-to-end model for\nFused Acoustic and Text Speech Translation (FAT-ST). Experiments on three\ntranslation directions show that by fine-tuning from FAT-MLM, our proposed\nspeech translation models substantially improve translation quality by up to\n+5.9 BLEU.\n