Patent №
US 7,899,665
Granted
—
Owner
—
Lab
—
AI components
4
ml · nlp · kr · hardware
Assignment
None on record
Dataset
AIPD
2023_r1 edition
Application
10922100
Embodiments of the present invention can gather data from native language sources to produce a valid collation sequence that is appropriate for a particular language and application. Sequences of characters in this data are tested to determine strength levels used by the given language. The data is also recursively probed with other sequences to test for contractions and identify expansions. Sequences in the data may then be compared against a known or predetermined sequence to generate a set of sorting rules that is specific to the language and application. The rules are formatted to replicate the sorting order found in the data.