Image-to-image translation (i2i) networks suffer from entanglement effects in\npresence of physics-related phenomena in target domain (such as occlusions,\nfog, etc), lowering altogether the translation quality, controllability and\nvariability. In this paper, we propose a general framework to disentangle\nvisual traits in target images. Primarily, we build upon collection of simple\nphysics models, guiding the disentanglement with a physical model that renders\nsome of the target traits, and learning the remaining ones. Because physics\nallows explicit and interpretable outputs, our physical models (optimally\nregressed on target) allows generating unseen scenarios in a controllable\nmanner. Secondarily, we show the versatility of our framework to neural-guided\ndisentanglement where a generative network is used in place of a physical model\nin case the latter is not directly accessible. Altogether, we introduce three\nstrategies of disentanglement being guided from either a fully differentiable\nphysics model, a (partially) non-differentiable physics model, or a neural\nnetwork. The results show our disentanglement strategies dramatically increase\nperformances qualitatively and quantitatively in several challenging scenarios\nfor image translation.\n
Paper
References (100)
Scroll for more · 38 remaining