Parts of Speech-Grounded Subspaces in Vision-Language Models

Latent image representations arising from vision-language models have proved immensely useful for a variety of downstream tasks. However, their utility is limited by their entanglement with respect to different visual attributes. For instance, recent work has shown that CLIP image representations are often biased toward specific visual properties (such as objects or actions) in an unpredictable manner. In this paper, we propose to separate representations of the different visual modalities in CLIP's joint vision-language space by leveraging the association between parts of speech and specific visual modes of variation (e.g. nouns relate to objects, adjectives describe appearance). This is achieved by formulating an appropriate component analysis model that learns subspaces capturing variability corresponding to a specific part of speech, while jointly minimising variability to the rest. Such a subspace yields disentangled representations of the different visual properties of an image or text in closed form while respecting the underlying geometry of the manifold on which the representations lie. What's more, we show the proposed model additionally facilitates learning subspaces corresponding to specific visual appearances (e.g. artists' painting styles), which enables the selective removal of entire visual themes from CLIP-based text-to-image synthesis. We validate the model both qualitatively, by visualising the subspace projections with a text-to-image model and by preventing the imitation of artists' styles, and quantitatively, through class invariance metrics and improvements to baseline zero-shot classification.

Paper

Similar papers

Peer review

Reviewer yh2S7/10 · confidence 5/52023-07-01

Summary

The paper presents an innovative solution to address the problem of polysemy within CLIP's embedding space. The authors propose a novel approach that involves decomposing CLIP embeddings into distinct subspaces, with each subspace representing a specific part of speech. This decomposition technique enables the isolation of different parts of speech within a given sentence, thereby facilitating subsequent manipulation in downstream tasks. The experimental results presented in the paper demonstrate the efficacy of the proposed approach in effectively eliminating properties associated with particular parts of speech during CLIP's text-to-image generation process.

Strengths

The paper presents a novel approach to decompose the embedding space of CLIP. Theoretical analysis and experimental results provide compelling evidence that the proposed approach effectively disentangles properties related to different parts of speech within the embedding space. This work is of significant value to researchers, as comprehending and manipulating the embedding space learned by deep neural networks is both crucial and challenging. Understanding the features that embeddings can represent and learning how to manipulate them is essential to the improvement of DNNs. The paper is written clearly and is well-structured.

Weaknesses

While the paper's exploration of subspace decomposition focuses on addressing the polysemy issues associated with part of speech in CLIP embeddings, it is important to note that there are instances of polysemy that cannot be disambiguated solely based on part of speech. For example, consider the word "crane," which can refer to both a bird and a machine. These instances present case-by-case ambiguities, and it remains unclear whether the proposed method can be extended to tackle such scenarios successfully as there are no universal subspaces that can disentangle all of them.

Questions

In the context of addressing polysemy and disambiguation, would it be more straightforward to incorporate more detailed descriptions in the prompts? Could you please elaborate on the advantages of the proposed method over using prompts with additional details?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

Yes

Reviewer Qiru5/10 · confidence 2/52023-07-02

Summary

The paper gives a closed-form solution to project CLIP representation of image/text into a subspace with disentangled modes. The proposed method is demonstrated qualitatively in text-to-image generation, and quantitatively by zero-shot classification.

Strengths

- The subspace projection proposed by the paper with a closed-form solution can directly be applied to models based on CLIP, without further training. - Qualitative results demonstrate the effectiveness in some extent.

Weaknesses

- The motivation of this work is CLIP's representation is biased and unpredictable. The paper proposes to learn sub-space representations for content and appearance. There are many diffusion-based works on guided generation for controlling contents and appearances. However, the paper (claims to have wide applications in generation tasks) fails to compare to any current work on this topic. - Limited quantitative studies, except for zero-shot classification. In summary, the close-form projection proposed in this paper is simple/fast and (qualitatively) effective in the some examples shown. But the paper lacks comparisons to recent works on diffusion generation with controllable contents and appearances.

Questions

N/A

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

2 fair

Presentation

2 fair

Contribution

2 fair

Limitations

The limitation is discussed, but only for the dimensionality of sub-space (a hyper-parameter k), which is superficial.

Reviewer QWgZ7/10 · confidence 4/52023-07-05

Summary

The following work proposes a geometry-aware approach to identifying subspace projections within the CLIP embedding space. The projections allow one to limit the CLIP embeddings to the subspaces corresponding to individual parts of speech (noun, adj... ). This allows for more fine-grained controllability when using CLIP embeddings for downstream tasks such as text to image synthesis. Notably, the authors take into account the non-euclidean nature of CLIP embeddings (located on the surface of a hypersphere) by mapping it a specific tangent space first. Experiments demonstrate additional controllability with regards to visual style when given access to a part-of-speech partitioning of the CLIP embedding space.

Strengths

- Principled approach to handling the non-euclidean nature of the CLIP embedding space. Personally I think this is an important topic that is often overlooked in many applied works building on top of CLIP and the recent improved VQGAN architecture, both of which employ a normalized embedding space. While the geometry-aware formulation is perhaps not the most novel contribution of this work, I believe it may serve as an important blueprint for future works relying on normalized embeddings. - Overall, the closed-form linear formulation of the subspace solving objective provides a low complexity but effective solution to the problem statement.

Weaknesses

- As much as the geometry aware formulation is mathematically justified, it would be nice to see some sort of experiment that demonstrates a significant loss in downstream task performance or even a simple plot based analysis as in the supplementary if one were to ignore it. - The visualization for lambda selection in supplementary figure 11 is unclear. The plotting software clearly overlays each point cloud on top of the others. As such, it is difficult to visually confirm the spread of all the point clouds except for the last one rendered, which I believe is the yellow adverb cloud. In order to properly visualize this, I think we would have to have juxtaposed separate plots for each cloud with the same axis scaling and offset. - While I am aware that there are many standards for placement of related works, I think it would be best for an earlier placement before methods given the proposed formulation's close relationship to fisher discriminant analysis.

Questions

See weaknesses.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

4 excellent

Contribution

4 excellent

Limitations

Adequately addressed.

Reviewer 3FWH7/10 · confidence 5/52023-07-07

Summary

This work proposes to learn subspaces for disentangling the visual representation in the CLIP space, based on the parts of speech of the prompt. A closed form solution is presented where the norm corresponding to the embedding of word of interest has a maximum norm while the norm of the rest is minimized. Qualitative results are shown on the CLIP-based TTIM from LAION where visual results are presented by killing on of the subspaces (noun or adjective) and quantitative through performance on a class invariance metric and through zero-shot classification. The method performs better than the prior art.

Strengths

+ The work is very well written, clearly motivated and well presented. + This is the first work which attempts to disentangle the subspaces using POS in CLIP based embedding models. + The method is easy to implement with the closed-form solution. + Qualitative and quantitative evaluation are performed to show the utility of POS guided subspace projection.

Weaknesses

- Effect of the prompts: From the quantitative results the role of the subspaces is not clear. For example, by removing noun from "Van Gogh" it generates the painting. Painting is also a noun. Therefore, the distinction is not clear. - In the qualitative examples (Figure 5), the images also change and the semantics on removing one POS subspace do not guarantee that the original image's style is preserved. - It would be good to show the results with multiple samples from the given prompt on the original dataset to show if the method actually works or just picks up on the partial prompts which would also work with the baseline model. For example, "A mutlicolored Penguin" and "A penguin" are very general prompts and they can have similar results without the subspace projection. - "Disentanglement" in multimodal approaches has been presented in prior work [1,2]. 1. Fast, Diverse and Accurate Image Captioning Guided By Part-of-Speech. CVPR 2019. 2. Diverse image captioning with context-object split latent spaces.

Questions

1. How would different partial prompts in the original baseline model compare with the proposed approach as pointed to in weakness? 2. Are there insights into persevering the underlying style of the image generated from the prompt? Example removing only snow from the already generated "snowy" images of NYC?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

There is no potential negative societal impact of their work.

Reviewer 3FWH2023-08-11

Thank you for the rebuttal

I have read the rebuttal which addresses all the points from the reviewer comments. The paper makes good contributions for controllable generation. I have updated my score accordingly.

Reviewer QWgZ2023-08-12

Concerns appropriately addressed

I have read the authors' responses to all reviews and am satisfied with their responses. As such, I retain my original rating of "Accept".

Reviewer yh2S2023-08-14

The authors' responses have addressed my questions. I've updated my rating to accept.

Reviewer Qiru2023-08-15

Thanks for the response. I've now updated my score accordingly.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC