ADAPTIVE OCR FOR BOOKS

Patent №

US 7,480,411

Granted

2009-01-20

Filed 2008

Owner

INTERNATIONAL BUSINESS MACHINE CORPORATION

Lab

AI components

3

ml · nlp · vision

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

12040946

A system/method is presented for scanning entire books or document all at once using an adaptive process where the book or document has known fonts and unknown fonts. The known fonts are processed through a verification system where sure words and error words are determined. Both the sure words and error words are sent to OCR training where they are re-OCR'ed and repeatedly verified until they meet a predetermined quality criteria. Characters or word not meeting the predetermined quality criteria receive additional OCR training until all the characters and words pass the predetermined quality criteria. Unknown fonts are scanned and clustered together by shape. Outliers in the shapes are manually key-in. Those symbols that are manually classified go to OCR training and then to the known type optimization process.

AI classification

Vision1.00
Natural language1.00
Machine learning1.00
AI hardware0.27
Speech0.06
Knowledge representation0.03
Evolutionary computation0.01
Planning0.00

Ownership

INTERNATIONAL BUSINESS MACHINE CORPORATION

assignment · 205870582

Assignors

TZADOK, ASAF, WALACH, EUGENIUSZ

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC