LOW-LATENCY MULTI-SPEAKER SPEECH RECOGNITION

Patent №

US 11,475,898

Granted

2022-10-18

Filed 2019

Owner

APPLE INC.

Lab

AI components

5

ml · nlp · speech · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16534902

Systems and processes for operating an intelligent automated assistant are provided. In one example, a method includes receiving mixed speech data representing utterances of a target speaker and utterances of one or more interfering audio sources. The method further includes obtaining a target speaker representation, which represents speech characteristics of the target speaker; and determining, using a learning network, probability distributions of phonetic elements directly from the mixed speech data. The inputs of the learning network include the mixed speech data and the target speaker representation. An output of the learning network includes the probability distributions of phonetic elements. The method further includes generating text corresponding to the utterances of the target speaker based on the probability distributions of the phonetic elements; and providing a response to the target speaker based on the text corresponding to the utterances of the target speaker.

Machine learningNatural languageSpeechPlanningAI hardwareG10L 17/00G10L 15/20G10L 17/02G10L 17/04G10L 17/18G10L 21/0272

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
AI hardware0.97
Planning0.74
Knowledge representation0.42
Vision0.04
Evolutionary computation0.00

Ownership

APPLE INC.

assignment · 514130786

Assignors

DELFARAH, MASOOD, ABDELHAMID, OSSAMA A., HWANG, KYUYEON, MCALLASTER, DONALD R., SINISCALCHI, SABATO MARCO

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC