The ETAPE corpus for the evaluation of speech-based TV content processing in the French language

The paper presents a comprehensive overview of existing data for the evaluation of spoken content processing in a multimedia framework for the French language. We focus on the ETAPE corpus which will be made publicly available by ELDA mid 2012, after completion of the evaluation campaign, and recall existing resources resulting from previous evaluation campaigns. The ETAPE corpus consists of 30 hours of TV and radio broadcasts, selected to cover a wide variety of topics and speaking styles, emphasizing spontaneous speech and multiple speaker areas.

Paper

Full text

PDF

The ETAPE corpus for the evaluation of speech-based TV content processing in the French language

Semantic Scholar · Computer Science · 2012

Abstract

The paper presents a comprehensive overview of existing data for the evaluation of spoken content processing in a multimedia framework for the French language. We focus on the ETAPE corpus which will be made publicly available by ELDA mid 2012, after completion of the evaluation campaign, and recall existing resources resulting from previous evaluation campaigns. The ETAPE corpus consists of 30 hours of TV and radio broadcasts, selected to cover a wide variety of topics and speaking styles, emphasizing spontaneous speech and multiple speaker areas.

References (16)

07Entit´es nomm´ees structur´ees : guide d’annotation quaero. Technical Report 2011-042011
09Anal - yse syntaxique du franais : des constituants aux dpen - dances . In TALN . Délégation Général l ’ Armement , 2008 . ESTER 2 : convention d ’ annotation détaillée et enrichie2009

Scroll for more · 4 remaining

Similar papers

© 2026 NYSGPT2525 LLC