ADAPTIVE OCR FOR BOOKS

Patent №

US 7,627,177

Granted

2009-12-01

Filed 2008

Owner

INTERNATIONAL BUSINESS MACHINE CORPORATION

Lab

AI components

5

ml · nlp · vision · speech · kr

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

12276907

A system is presented for scanning entire books or document all at once using an adaptive process where the book or document has known fonts and unknown fonts. The known fonts are processed through a verification system where sure words and error words are determined. Both the sure words and error words are sent to OCR training where they are re-OCR'ed and repeatedly verified until they meet a predetermined quality criteria. Characters or words not meeting the predetermined quality criteria receive additional OCR training until all the characters and words pass the predetermined quality criteria. Unknown fonts are scanned and clustered together by shape. Outliers in the shapes are manually keyed-in. Those symbols that are manually classified go to OCR training and then to the known type optimization process.

AI classification

Vision1.00
Machine learning1.00
Natural language1.00
Speech0.98
Knowledge representation0.68
AI hardware0.18
Evolutionary computation0.07
Planning0.00

Ownership

INTERNATIONAL BUSINESS MACHINE CORPORATION

assignment · 218830179

Assignors

TZADOK, ASAF, WALACH, EUGENIUSZ

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC