Qimera: Data-free Quantization with Synthetic Boundary Supporting Samples

Model quantization is known as a promising method to compress deep neural\nnetworks, especially for inferences on lightweight mobile or edge devices.\nHowever, model quantization usually requires access to the original training\ndata to maintain the accuracy of the full-precision models, which is often\ninfeasible in real-world scenarios for security and privacy issues. A popular\napproach to perform quantization without access to the original data is to use\nsynthetically generated samples, based on batch-normalization statistics or\nadversarial learning. However, the drawback of such approaches is that they\nprimarily rely on random noise input to the generator to attain diversity of\nthe synthetic samples. We find that this is often insufficient to capture the\ndistribution of the original data, especially around the decision boundaries.\nTo this end, we propose Qimera, a method that uses superposed latent embeddings\nto generate synthetic boundary supporting samples. For the superposed\nembeddings to better reflect the original distribution, we also propose using\nan additional disentanglement mapping layer and extracting information from the\nfull-precision model. The experimental results show that Qimera achieves\nstate-of-the-art performances for various settings on data-free quantization.\nCode is available at https://github.com/iamkanghyunchoi/qimera.\n

Paper

References (54)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC