A Structured Three-Stage Pipeline for Compositional Text-to-Image Generation with Editable Layouts and Object-Wise Attention

When the prompt is prompting several objects in the scene, text-to-image models still fail because they often mismatch attributes, place objects incorrectly, or are left without important relationships. That makes them inadequate for real-world design, visualization, or creative work where users require precise, predictable control. Existing strategies, such as prompt tricks, attention tweaks, or layout-guided diffusion help in some ways, but they tend to fail: they not ground attributes properly, they deteriorate on complex scenes, or they become challenging to edit after generation has begun. In order to overcome these limitations, we provide a three-stage compositional synthesis pipeline for fine-grained, variable image generation. First, a language model looks at the prompt in terms of clean, sequenced presentation of entities, attributes, and spatial relation. Next, the system uses this structure to predict an editable intermediate layout and an attribute map for each object. Finally, the image is generated by diffusion model accompanied by attention-masking system controlling the level of objects during synthesis. The user can then edit layouts before rendering, ensuring that the final output matches the prompt and the structure. Our approach consistently improves object count accuracy, attribute accuracy and spatial coherence within experiments, and at the same time without compromising visual quality.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC