Summary
i)This paper uncovers the impact of representational inductive biases on Latent Diffusion Models through one-shot tasks, particularly in the realm of human-like sketching. It compares three distinct groups of regularizers: a standard baseline, supervised methods, and a third group consisting of self-supervised techniques.
ii)This paper aims to uncover the strategies employed by each regularization method to generalize to novel categories.
iii)This paper conducts a comprehensive comparative analysis of the effectiveness of different regularization methods in one-shot drawing tasks. The study explores various dimensions, examining the performance and differences of various regularization strategies in such tasks.
Strengths
This paper is dedicated to analyzing the specific manifestations of various inductive biases in one-shot drawing tasks, contributing new findings and demonstrating a high level of originality. The paper is well written, with a clear and logical structure, reflecting a rigorous academic attitude. The experimental section is well-designed, with substantial research effort, and the execution process strictly adheres to scientific methods. The comparative analysis strategy used is both comprehensive and detailed, effectively ensuring the precision and credibility of the research findings. Additionally, the conclusions of this paper provide valuable insights for the field of one-shot drawing tasks and have significant guiding significance for subsequent research.
Weaknesses
i)This paper conducted an extensive analysis of the impact of different inductive biases in one-shot drawing tasks. However, the analysis remains superficial, merely briefly revealing the experimental results of various methods without delving into the fundamental reasons behind the differences in outcomes among the methods. Furthermore, the paper does not propose specific solutions to the research questions addressed.
ii)The experimental method adopted in this study is not limited to specific modalities and is not restricted by data scale. This method may be applicable to the generation tasks of more modalities(such as photos or text2img). However, the experiment specifically chose handwriting and sketches as the research subjects for one-shot generation tasks.
iii) The core argument proposed by this paper is the generation of samples that resemble human-drawn sketches or handwriting. However, the paper does not provide sufficient elaboration on how to quantify the “human-like” characteristics of the samples, especially with an in-depth analysis from the perspective of stroke features.
Questions
Major:
i)In one-shot drawing tasks, originality constitutes a key evaluation metric. However, when measuring the “human-like” characteristics of samples, the impact of originality is relatively minor, and its role in the evaluation process appears to be more limited.
ii)
Does this paper quantitatively analyze the stroke correlation between the generated sketches and those drawn by humans? Although the generalization curve has been considered in the evaluation of originality and recognizability, the paper does not explicitly reveal the interrelationship between the strokes.
iii)Has this study explored the reasons for or proposed hypotheses about the performance differences exhibited by various inductive bias methods in one-shot drawing tasks?
Minor:
i)Given that classification models may be influenced by their own biases or uneven distributions in the training data, the recognizability of the model does not necessarily equate to the “human-like” level of the samples. Does the paper take this potential issue into consideration?
ii)In the exploration of effective methods to improve the performance of one-shot drawing tasks, did this paper consider approaches other than simply adding the prototype-based and Barlow regularizers with weights?
iii)
In Figure 2, does the “pixel baseline” refer to the use of rasterized sketches during training or testing? On this point, this paper could analyze the differential impact of inductive biases when dealing with vector versus raster sketches.
Limitations
No. How the evaluation criteria effectively align with human judgment or aesthetic standards is a question of considerable research value, especially in generative tasks. Particularly when dealing with sketches or handwriting, the consideration of strokes is an indispensable key dimension that cannot be overlooked.