Improving Visual Quality of Image Synthesis by A Token-based Generator with Transformers

We present a new perspective of achieving image synthesis by viewing this\ntask as a visual token generation problem. Different from existing paradigms\nthat directly synthesize a full image from a single input (e.g., a latent\ncode), the new formulation enables a flexible local manipulation for different\nimage regions, which makes it possible to learn content-aware and fine-grained\nstyle control for image synthesis. Specifically, it takes as input a sequence\nof latent tokens to predict the visual tokens for synthesizing an image. Under\nthis perspective, we propose a token-based generator (i.e.,TokenGAN).\nParticularly, the TokenGAN inputs two semantically different visual tokens,\ni.e., the learned constant content tokens and the style tokens from the latent\nspace. Given a sequence of style tokens, the TokenGAN is able to control the\nimage synthesis by assigning the styles to the content tokens by attention\nmechanism with a Transformer. We conduct extensive experiments and show that\nthe proposed TokenGAN has achieved state-of-the-art results on several\nwidely-used image synthesis benchmarks, including FFHQ and LSUN CHURCH with\ndifferent resolutions. In particular, the generator is able to synthesize\nhigh-fidelity images with 1024x1024 size, dispensing with convolutions\nentirely.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC