Visual Features for Context-Aware Speech Recognition

Automatic transcriptions of consumer generated multi-media content such as “Youtube” videos still exhibit high word error rates. Such data typically occupies a very broad domain, has been recorded in challenging conditions, with cheap hardware and a focus on the visual modality, and may have been post-processed or edited.

Paper

Similar papers

© 2026 NYSGPT2525 LLC