Meta-review
The paper introduces a unified transformer model Show-o that can do multimodal understanding and generation in one shared transformer model. The model combines autoregressive and discrete diffusion modeling to handle different input and output modalities. This approach with a 1.3B base model achieves decent results on various tasks such as visual question-answering and text-to-image generation.
This paper presents a new approach of combining autoregressive and diffusion modeling, which is an insightful and valuable exploration in the direction of unified multimodal learning.
The reviewers are all positive about this paper.
I recommend accepting this paper.
Additional comments on reviewer discussion
Reviewer 4YnC raises concerns about insufficient comparison, special token design, and ability to manage interleaved input and output
The score is maintained as 6 after rebuttal.
Reviewer vpwM raises concerns about high training cost, inference speed, issues about scaling up, conflicts between tasks, and more benchmarks.
The score is increased from 6 to 8 after rebuttal.
Reviewer 3aSb raises concerns about architecture choice, unfair comparison, the generation ability with continuous tokens, impact of classifier-free guidance.
The score is maintained as 6 after rebuttal.
Reviewer 8ezu raises concerns about clarify over prior discrete diffusion modeling, continuous tokens for generation, inference speed.
The score is increased from 5 to 6 after rebuttal.
The ratings are mixed in the initial reviews. (6, 6, 6, 5), and becomes positive after the rebuttal. (6, 8, 6, 6)
The concerns are addressed by authors in rebuttal via more experiments, comparison and analysis.