Patent №
US 10,679,643
Granted
—
Owner
—
Lab
—
AI components
4
ml · nlp · speech · hardware
Assignment
None on record
Dataset
AIPD
2023_r1 edition
Application
15691546
A method, computer readable medium, and system are disclosed for audio captioning. A raw audio waveform including a non-speech sound is received and relevant features are extracted from the raw audio waveform using a recurrent neural network (RNN) acoustic model. A discrete sequence of characters represented in a natural language is generated based on the relevant features, where the discrete sequence of characters comprises a caption that describes the non-speech sound.