REDUCING DIGEST STORAGE CONSUMPTION BY TRACKING SIMILARITY ELEMENTS IN A DATA DEDUPLICATION SYSTEM
Patent №
US 9,116,941
Granted
2015-08-25
Filed 2013
Owner
INTERNATIONAL BUSINESS MACHINES CORPORATION
Lab
AI components
3
ml · nlp · kr
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
13840438
For reducing digests storage consumption in a data deduplication system using a processor device in a computing environment, input data is partitioned into chunks, and the chunks are grouped into chunk sets. Digests are calculated for input data and stored in sets corresponding to the chunk sets. Similarity elements are calculated for the input data and the similarity elements are stored in a similarity search structure. The number of similarity elements associated with a chunk set which are currently contained in the similarity search structure is maintained for each chunk set, and when this number of a specific chunk set becomes lower than a threshold, the digests set associated with that chunk set are removed from the repository.
AI classification
Ownership
INTERNATIONAL BUSINESS MACHINES CORPORATION
assignment · 300200571
Assignors
ARONOVICH, LIOR
On an employer assignment, the assignors are typically the inventors.