SYSTEM AND METHOD FOR DISAMBIGUATING NON DIACRITIZED ARABIC WORDS IN A TEXT

Patent №

US 8,041,559

Granted

2011-10-18

Filed 2005

Owner

MACHINES CORPORATION, INTERNATIONAL BUSINESS

Lab

AI components

7

ml · nlp · vision · speech · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11299220

The present invention proposes a solution to the problem of word lexical disambiguation in Arabic texts. This solution is based on text domain-specific knowledge, which facilitates the automatic vowel restoration of modern standard Arabic scripts. Texts similar in their contents, restricted to a specific field or sharing a common knowledge can be grouped in a specific category or in a specific domain (examples of specific domains; sport, art, economic, science . . . ). The present invention discloses a method, system and computer program for lexically disambiguating non diacritized Arabic words in a text based on a learning approach that exploits; Arabic lexical look-up, and Arabic morphological analysis, to train the system on a corpus of diacritized Arabic text pertaining to a specific domain. Thereby, the contextual relationships of the words related to a specific domain are identified, based on the valid assumption that there is less lexical variability in the use of the words and their morphological variants within a domain compared to an unrestricted text.

AI classification

Natural language1.00
Speech1.00
AI hardware1.00
Machine learning1.00
Knowledge representation1.00
Planning0.85
Vision0.77
Evolutionary computation0.00

Ownership

MACHINES CORPORATION, INTERNATIONAL BUSINESS

assignment · 171700914

Assignors

EL-SHISHINY, HISHAM

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC