SEGMENTATION OF AUDIO DATA FOR INDEXING OF CONVERSATIONAL SPEECH FOR REAL-TIME OR POSTPROCESSING APPLICATIONS

Patent №

US 5,655,058

Granted

1997-08-05

Filed 1994

Owner

XEROX CORPORATION

Lab

AI components

4

ml · nlp · vision · speech

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

08226519

A method for segmenting audio data, comprising speech from a plurality of individual speakers, according to speaker is provided. The method comprises providing individual HMMs for each individual speaker, each individual HMM including at least one state, and constructing a speaker network HMM by connecting the individual HMMs in parallel. The audio data is then divided into segments by determining a most likely sequence of states through the speaker network HMM, each of the segments being associated with one of the individual HMMs. Afterward, the speaker of each of the segments is identified. The segmented data may be used to form an index into the audio data according to speaker.

AI classification

Speech1.00
Natural language1.00
Machine learning1.00
Vision0.84
AI hardware0.02
Knowledge representation0.00
Evolutionary computation0.00
Planning0.00

Ownership

XEROX CORPORATION

assignment · 69560789

Assignors

BALASUBRAMANIAN, VIJAY, CHEN, FRANCINE R., CHOU, PHILIP A., KIMBER, DONALD G., POON, ALEX D., WEBER, KARON A.

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC