RealCustom++: Representing Images as Real Textual Word for Real-Time Customization

Text-to-image customization aims to generate images that align with both the given text and the subject in the given image. Existing works follow the pseudo-word paradigm, which represents the subject as a non-existent pseudo word and combines it with other text to generate images. However, the pseudo word inherently conflicts and entangles with other real words, resulting in a dual-optimum paradox between the subject similarity and text controllability. To address this, we propose RealCustom++, a novel real-word paradigm that represents the subject with a non-conflicting real word to generate a coherent guidance image and corresponding subject mask, there by disentangling the influence scopes of the text and subject for simultaneous optimization. Specifically, RealCustom++ introduces a train-inference decoupled framework: (1) during training, it learns a general alignment between visual conditions and all real text words; and (2) during inference, a dual-branch architecture is employed, where the Guidance Branch produces the subject guidance mask, and the Generation Branch utilizes this mask to customize the generation of the specific real word exclusively within subject-relevant regions. Extensive experiments validate RealCustom++s superior performance, which improves controllability by 7.48%, similarity by 3.04% and quality by 76.43% simultaneously. Moreover, RealCustom++ further improves controllability by 4.6% and multi-subject similarity by 6.34% for multisubject customization

Paper

Similar papers

© 2026 NYSGPT2525 LLC