REDUCING DIGEST STORAGE CONSUMPTION BY TRACKING SIMILARITY ELEMENTS IN A DATA DEDUPLICATION SYSTEM

Patent №

US 9,116,941

Granted

2015-08-25

Filed 2013

Owner

INTERNATIONAL BUSINESS MACHINES CORPORATION

AI components

3

ml · nlp · kr

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

13840438

For reducing digests storage consumption in a data deduplication system using a processor device in a computing environment, input data is partitioned into chunks, and the chunks are grouped into chunk sets. Digests are calculated for input data and stored in sets corresponding to the chunk sets. Similarity elements are calculated for the input data and the similarity elements are stored in a similarity search structure. The number of similarity elements associated with a chunk set which are currently contained in the similarity search structure is maintained for each chunk set, and when this number of a specific chunk set becomes lower than a threshold, the digests set associated with that chunk set are removed from the repository.

Machine learningNatural languageKnowledge representationG06F 16/2365G06F 3/0641G06F 16/1752G06F 16/215G06F 16/278G06F 16/285G06F 16/955

AI classification

Knowledge representation1.00
Machine learning0.99
Natural language0.53
Planning0.15
AI hardware0.14
Vision0.00
Evolutionary computation0.00
Speech0.00

Ownership

INTERNATIONAL BUSINESS MACHINES CORPORATION

assignment · 300200571

Assignors

ARONOVICH, LIOR

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC