In egocentric videos, actions occur in quick succession. We capitalise on the\naction's temporal context and propose a method that learns to attend to\nsurrounding actions in order to improve recognition performance. To incorporate\nthe temporal context, we propose a transformer-based multimodal model that\ningests video and audio as input modalities, with an explicit language model\nproviding action sequence context to enhance the predictions. We test our\napproach on EPIC-KITCHENS and EGTEA datasets reporting state-of-the-art\nperformance. Our ablations showcase the advantage of utilising temporal context\nas well as incorporating audio input modality and language model to rescore\npredictions. Code and models at: https://github.com/ekazakos/MTCN.\n