CONSTRUCTING A TRANSLATION LEXICON FROM COMPARABLE, NON-PARALLEL CORPORA

Patent №

US 7,620,538

Granted

2009-11-17

Filed 2003

Owner

SOUTHERN CALIFORNIA, UNIVERSITY OF

Lab

AI components

5

ml · nlp · speech · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

10401124

A machine translation system may use non-parallel monolingual corpora to generate a translation lexicon. The system may identify identically spelled words in the two corpora, and use them as a seed lexicon. The system may use various clues, e.g., context and frequency, to identify and score other possible translation pairs, using the seed lexicon as a basis. An alternative system may use a small bilingual lexicon in addition to non-parallel corpora to learn translations of unknown words and to generate a parallel corpus.

AI classification

Natural language1.00
Machine learning1.00
Speech1.00
Knowledge representation1.00
AI hardware1.00
Planning0.36
Vision0.30
Evolutionary computation0.00

Ownership

SOUTHERN CALIFORNIA, UNIVERSITY OF

assignment · 137550972

Assignors

MARCU, DANIEL, KNIGHT, KEVIN, MUNTEANU, DRAGOS STEFAN, KOEHN, PHILIPP

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC