NowYouSee Me: Context-Aware Automatic Audio Description

Audio Description (AD) plays a pivotal role as an application system aimed at\nguaranteeing accessibility in multimedia content, which provides additional\nnarrations at suitable intervals to describe visual elements, catering\nspecifically to the needs of visually impaired audiences. In this paper, we\nintroduce $\\mathrm{CA^3D}$, the pioneering unified Context-Aware Automatic\nAudio Description system that provides AD event scripts with precise locations\nin the long cinematic content. Specifically, $\\mathrm{CA^3D}$ system consists\nof: 1) a Temporal Feature Enhancement Module to efficiently capture longer term\ndependencies, 2) an anchor-based AD event detector with feature suppression\nmodule that localizes the AD events and extracts discriminative feature for AD\ngeneration, and 3) a self-refinement module that leverages the generated output\nto tweak AD event boundaries from coarse to fine. Unlike conventional methods\nwhich rely on metadata and ground truth AD timestamp for AD detection and\ngeneration tasks, the proposed $\\mathrm{CA^3D}$ is the first end-to-end\ntrainable system that only uses visual cue. Extensive experiments demonstrate\nthat the proposed $\\mathrm{CA^3D}$ improves existing architectures for both AD\nevent detection and script generation metrics, establishing the new\nstate-of-the-art performances in the AD automation.\n

Paper

References (51)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC