METHOD FOR SEARCHING NON-TOKENIZED TEXT FOR MATCHES AGAINST A KEYWORD DATA STRUCTURE
Patent №
US 6,263,333
Granted
2001-07-17
Filed 1998
Owner
INTERNATIONAL BUSINESS MACHINES CORPORATION
Lab
AI components
5
ml · nlp · speech · kr · planning
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
09177034
A method for searching a non-tokenized text string for matches against a keyword data structure organized as a set of one or more keyword objects. The method begins by (a) indexing into the keyword data structure using a character in the non-tokenized text string. Preferably, the character is a Unicode value. The routine then continues by (b) comparing a portion of the non-tokenized text string to a keyword object. If the portion of the non-tokenized text string matches the keyword object, the routine saves the keyword object in a match list. If, however, the portion of the non-tokenized text string does not match the keyword object and there are no other keyword objects that share a root with the non-matched keyword object, the routine repeats step (a) with a new character. These steps are then repeated until all characters in the non-tokenized text string have been analyzed against the keyword data structure.
AI classification
Ownership
INTERNATIONAL BUSINESS MACHINES CORPORATION
assignment · 112030031
Assignors
HOUCHIN, ALICE M., WOOD, DOUGLAS A.
On an employer assignment, the assignors are typically the inventors.