While neural networks with attention mechanisms have achieved superior\nperformance on many natural language processing tasks, it remains unclear to\nwhich extent learned attention resembles human visual attention. In this paper,\nwe propose a new method that leverages eye-tracking data to investigate the\nrelationship between human visual attention and neural attention in machine\nreading comprehension. To this end, we introduce a novel 23 participant eye\ntracking dataset - MQA-RC, in which participants read movie plots and answered\npre-defined questions. We compare state of the art networks based on long\nshort-term memory (LSTM), convolutional neural models (CNN) and XLNet\nTransformer architectures. We find that higher similarity to human attention\nand performance significantly correlates to the LSTM and CNN models. However,\nwe show this relationship does not hold true for the XLNet models -- despite\nthe fact that the XLNet performs best on this challenging task. Our results\nsuggest that different architectures seem to learn rather different neural\nattention strategies and similarity of neural to human attention does not\nguarantee best performance.\n