BUILDING A TRANSLATION LEXICON FROM COMPARABLE, NON-PARALLEL CORPORA

Patent №

US 8,234,106

Granted

2012-07-31

Filed 2009

Owner

LANGUAGE WEAVER, INC.

+1 more

Lab

AI components

6

ml · nlp · speech · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

12576110

A machine translation system may use non-parallel monolingual corpora to generate a translation lexicon. The system may identify identically spelled words in the two corpora, and use them as a seed lexicon. The system may use various clues, e.g., context and frequency, to identify and score other possible translation pairs, using the seed lexicon as a basis. An alternative system may use a small bilingual lexicon in addition to non-parallel corpora to learn translations of unknown words and to generate a parallel corpus.

AI classification

Natural language1.00
Machine learning1.00
Speech1.00
Knowledge representation0.99
AI hardware0.76
Planning0.57
Vision0.02
Evolutionary computation0.00

Ownership

LANGUAGE WEAVER, INC.

assignment · 241420888

UNIVERSITY OF SOUTHERN CALIFORNIA

assignment · 245690486

Assignors

MARCU, DANIEL, KNIGHT, KEVIN, MUNTEANU, DRAGOS S, KOEHN, PHILIPP

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC