Building a RAPPOR with the Unknown: Privacy-Preserving Learning of Associations and Data Dictionaries

Techniques based on randomized response enable the collection of potentially\nsensitive data from clients in a privacy-preserving manner with strong local\ndifferential privacy guarantees. One of the latest such technologies, RAPPOR,\nallows the marginal frequencies of an arbitrary set of strings to be estimated\nvia privacy-preserving crowdsourcing. However, this original estimation process\nrequires a known set of possible strings; in practice, this dictionary can\noften be extremely large and sometimes completely unknown.\n In this paper, we propose a novel decoding algorithm for the RAPPOR mechanism\nthat enables the estimation of "unknown unknowns," i.e., strings we do not even\nknow we should be estimating. To enable learning without explicit knowledge of\nthe dictionary, we develop methodology for estimating the joint distribution of\ntwo or more variables collected with RAPPOR. This is a critical step towards\nunderstanding relationships between multiple variables collected in a\nprivacy-preserving manner.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC