Systems and methods for identifying parallel documents and sentence fragments in multilingual document collections
Patent №
US 8,943,080
Granted
2015-01-27
Filed 2006
Owner
UNIVERSITY OF SOUTHERN CALIFORNIA
Lab
—
AI components
2
ml · nlp
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
11635248
Systems, computer programs, and methods for identifying parallel documents and/or fragments in a bilingual collection are provided. The method for identifying parallel sub-sentential fragments in a bilingual collection comprises translating a source document from a bilingual collection. The method further includes querying a target library associated with the bilingual collection using the translated source document, and identifying one or more target documents based on the query. Subsequently, a source sentence associated with the source document is aligned to one or more target sentences associated with the one or more target documents. Finally, the method includes determining whether a source fragment associated with the source sentence comprises a parallel translation of a target fragment associated with the one or more target sentences.
AI classification
Ownership
UNIVERSITY OF SOUTHERN CALIFORNIA
assignment · 186590551
Assignors
MARCU, DANIEL, MUNTEANU, DRAGOS STEFAN
On an employer assignment, the assignors are typically the inventors.