Audio Description (AD) plays a pivotal role as an application system aimed at\nguaranteeing accessibility in multimedia content, which provides additional\nnarrations at suitable intervals to describe visual elements, catering\nspecifically to the needs of visually impaired audiences. In this paper, we\nintroduce $\\mathrm{CA^3D}$, the pioneering unified Context-Aware Automatic\nAudio Description system that provides AD event scripts with precise locations\nin the long cinematic content. Specifically, $\\mathrm{CA^3D}$ system consists\nof: 1) a Temporal Feature Enhancement Module to efficiently capture longer term\ndependencies, 2) an anchor-based AD event detector with feature suppression\nmodule that localizes the AD events and extracts discriminative feature for AD\ngeneration, and 3) a self-refinement module that leverages the generated output\nto tweak AD event boundaries from coarse to fine. Unlike conventional methods\nwhich rely on metadata and ground truth AD timestamp for AD detection and\ngeneration tasks, the proposed $\\mathrm{CA^3D}$ is the first end-to-end\ntrainable system that only uses visual cue. Extensive experiments demonstrate\nthat the proposed $\\mathrm{CA^3D}$ improves existing architectures for both AD\nevent detection and script generation metrics, establishing the new\nstate-of-the-art performances in the AD automation.\n
Paper
References (51)
Scroll for more · 38 remaining