UNSUPERVISED SPEAKER SEGMENTATION OF MULTI-SPEAKER SPEECH DATA

Patent №

US 7,930,179

Granted

2011-04-19

Filed 2007

Owner

AT&T CORP.

Lab

AI components

5

ml · nlp · vision · speech · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11866125

Systems and methods for unsupervised segmentation of multi-speaker speech or audio data by speaker. A front-end analysis is applied to input speech data to obtain feature vectors. The speech data is initially segmented and then clustered into groups of segments that correspond to different speakers. The clusters are iteratively modeled and resegmented to obtain stable speaker segmentations. The overlap between segmentation sets is checked to ensure successful speaker segmentation. Overlapping segments are combined and remodeled and resegmented. Optionally, the speech data is processed to produce a segmentation lattice to maximize the overall segmentation likelihood.

AI classification

Machine learning1.00
Speech1.00
Natural language0.99
Vision0.96
AI hardware0.61
Knowledge representation0.07
Planning0.00
Evolutionary computation0.00

Ownership

AT&T CORP.

assignment · 381210977

Assignors

GORIN, ALLEN LOUIS, LIU, ZHU, PARTHASARATHY, SARANGARAJAN, ROSENBERG, AARON EDWARD

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC