"This is Houston. Say again, please". The Behavox system for the Apollo-11 Fearless Steps Challenge (phase II)

We describe the speech activity detection (SAD), speaker diarization (SD),\nand automatic speech recognition (ASR) experiments conducted by the Behavox\nteam for the Interspeech 2020 Fearless Steps Challenge (FSC-2). A relatively\nsmall amount of labeled data, a large variety of speakers and channel\ndistortions, specific lexicon and speaking style resulted in high error rates\non the systems which involved this data. In addition to approximately 36 hours\nof annotated NASA mission recordings, the organizers provided a much larger but\nunlabeled 19k hour Apollo-11 corpus that we also explore for semi-supervised\ntraining of ASR acoustic and language models, observing more than 17% relative\nword error rate improvement compared to training on the FSC-2 data only. We\nalso compare several SAD and SD systems to approach the most difficult tracks\nof the challenge (track 1 for diarization and ASR), where long 30-minute audio\nrecordings are provided for evaluation without segmentation or speaker\ninformation. For all systems, we report substantial performance improvements\ncompared to the FSC-2 baseline systems, and achieved a first-place ranking for\nSD and ASR and fourth-place for SAD in the challenge.\n

Paper

References (33)

Scroll for more · 21 remaining

Similar papers

© 2026 NYSGPT2525 LLC