Look and Listen: A Multi-modality Late Fusion Approach to Scene Classification for Autonomous Machines
The novelty of this study consists in a multi-modality approach to scene\nclassification, where image and audio complement each other in a process of\ndeep late fusion. The approach is demonstrated on a difficult classification\nproblem, consisting of two synchronised and balanced datasets of 16,000 data\nobjects, encompassing 4.4 hours of video of 8 environments with varying degrees\nof similarity. We first extract video frames and accompanying audio at one\nsecond intervals. The image and the audio datasets are first classified\nindependently, using a fine-tuned VGG16 and an evolutionary optimised deep\nneural network, with accuracies of 89.27% and 93.72%, respectively. This is\nfollowed by late fusion of the two neural networks to enable a higher order\nfunction, leading to accuracy of 96.81% in this multi-modality classifier with\nsynchronised video frames and audio clips. The tertiary neural network\nimplemented for late fusion outperforms classical state-of-the-art classifiers\nby around 3% when the two primary networks are considered as feature\ngenerators. We show that situations where a single-modality may be confused by\nanomalous data points are now corrected through an emerging higher order\nintegration. Prominent examples include a water feature in a city misclassified\nas a river by the audio classifier alone and a densely crowded street\nmisclassified as a forest by the image classifier alone. Both are examples\nwhich are correctly classified by our multi-modality approach.\n
Paper
References (23)
Scroll for more · 11 remaining