Mapping EEG Signals to Visual Stimuli: A Deep Learning Approach to Match vs. Mismatch Classification
Existing approaches to modeling associations between visual stimuli and brain responses are facing difficulties in handling between-subject variance and model generalization. Inspired by the recent progress in modeling speech-brain response, we propose a “match-vs-mismatch” deep learning model in this study to classify whether a video clip elicits neural responses in recorded EEG signals. Our model employs dilated convolutional neural networks and gated recurrent units to extract features from both EEG and video signals, enabling the learning of associations between visual content and corresponding neural recordings. We demonstrate that our proposed model achieves the highest accuracy on unseen subjects compared to other baseline models. Additionally, we assess inter-subject noise using a subject-level silhouette score in the embedding space, revealing that our model effectively mitigates inter-subject noise and significantly reduces the silhouette score. Furthermore, we investigate Grad-CAM activation scores, revealing that brain regions linked to language processing contribute most to model predictions, followed by regions associated with visual processing. These findings hold promise for advancing neural recording-based video reconstruction and related applications.
Paper
References (34)
Scroll for more · 22 remaining