CemiFace: Center-based Semi-hard Synthetic Face Generation for Face Recognition

Privacy issue is a main concern in developing face recognition techniques. Although synthetic face images can partially mitigate potential legal risks while maintaining effective face recognition (FR) performance, FR models trained by face images synthesized by existing generative approaches frequently suffer from performance degradation problems due to the insufficient discriminative quality of these synthesized samples. In this paper, we systematically investigate what contributes to solid face recognition model training, and reveal that face images with certain degree of similarities to their identity centers show great effectiveness in the performance of trained FR models. Inspired by this, we propose a novel diffusion-based approach (namely Center-based Semi-hard Synthetic Face Generation (CemiFace)) which produces facial samples with various levels of similarity to the subject center, thus allowing to generate face datasets containing effective discriminative samples for training face recognition. Experimental results show that with a modest degree of similarity, training on the generated dataset can produce competitive performance compared to previous generation methods.

Paper

Similar papers

Peer review

Reviewer Crof9/10 · confidence 5/52024-07-08

Summary

This paper proposes CemiFace, a novel diffusion-based approach for generating synthetic face images with varying levels of similarity to their identity centers. The authors argue that semi-hard negative samples, those with moderate similarity to the center, are crucial for training effective face recognition models. The core of CemiFace lies in its ability to control the similarity between generated images and the input (identity center) during the diffusion process. This is achieved by injecting a similarity controlling factor condition (m) that regulates the similarity level. The paper presents a comprehensive analysis of the relationship between sample similarity and face recognition performance, showing that semi-hard samples, generated with m close to 0, achieve the best accuracy. CemiFace demonstrates significant improvements over previous methods in terms of accuracy, particularly on pose-sensitive datasets. The paper further validates its effectiveness through qualitative visualizations and ablation studies that examine the impact of various factors, including training data, inquiry data, and the similarity controlling factor. Overall, this paper contributes a valuable approach to generating synthetic face datasets for face recognition with enhanced discriminative power. The method shows promise in mitigating privacy concerns associated with collecting and using real-world face data while maintaining robust recognition performance.

Strengths

Discovery of the importance of similarity control in synthetic face generation: CemiFace is motivated by the discovery that face images with certain degree of similarities to their identity centers show great effectiveness in the performance of trained FR models. This is an important discovery to the community of synthetic dataset generation. Unique use of similarity control: CemiFace introduces a similarity controlling factor (m) within the diffusion process, enabling the generation of faces with varying levels of similarity to the input image. This provides a fine-grained control over the generated data distribution, which is a unique feature compared to existing methods. Comprehensive analysis of similarity: The authors present a thorough analysis of the impact of different similarity levels on face recognition performance, validating their hypothesis about the importance of semi-hard samples. This analysis provides valuable insights into the relationship between data distribution and model effectiveness. Rigorous experimental evaluation: The paper conducts comprehensive experiments across various benchmark datasets and data volumes, comparing CemiFace with other state-of-the-art synthetic face generation methods. The ablation studies provide a detailed understanding of the influence of different parameters and factors on the model's performance. Robustness of CemiFace: The experiments demonstrate the robustness of CemiFace to different training data, inquiry data, and similarity controlling factors. The method consistently achieves superior results, demonstrating its effectiveness and generalizability.

Weaknesses

- An in-depth discussion on why face images with certain similarity is more beneficial as a training dataset for the face recognition model would strengthen the paper. For example, an analysis such as a similarity comparison the the real dataset and checking if the difficulty of the CemiFace synthetic dataset becomes closer to that of the real dataset would be nice. Other analysis that offers insights as to why certain similarity control is important would also be welcome.

Questions

- Written in weakness section.

Rating

9

Confidence

5

Soundness

4

Presentation

4

Contribution

4

Limitations

- This is a well written paper with a meaningful discovery on training with synthetic dataset. The paper would be better positioned in the venue NeurIPS if it would offer more insightful analysis on why similarity control is beneficial, on top of the empirical benefits.

Reviewer 5xKY5/10 · confidence 4/52024-07-09

Summary

The paper introduces an approach called CemiFace for generating synthetic face images to enhance face recognition (FR) models. The paper provides the first in-depth analysis of how FR model performance is influenced by samples with varying levels of similarity to the identity center, focusing particularly on center-based semi-hard samples. The authors propose a unique diffusion-based model that can generate face images with different levels of similarity to the identity center. This model can produce infinite center-based semi-hard face images for synthetic face recognition (SFR). The method can be extended to leverage large amounts of unlabeled data for training, providing an advantage over previous methods. Experimental results demonstrate that CemiFace significantly outperforms existing SFR methods, reducing the GAP-to-Real error by half and showcasing promising performance in synthetic face recognition.

Strengths

- Focusing on center-based semi-hard samples to enhance face recognition performance is a fresh problem formulation that addresses a notable gap in current methodologies. - The paper provides a solid experimental validation of its proposed approach. The authors investigate factors affecting performance degradation in synthetic face recognition and offer a hypothesis about the importance of mid-level similarity samples.

Weaknesses

- The method for determining GtR remains unclear. Justification regarding how the proposed model yields a low GtR is absent. Is this low GtR attributed to the utilization of real inquiry images? If so, what measures guarantee that the synthetic facial images remain uncorrelated with the real facial images? In other words, I have a reservation that the method may not generate "true" synthetic data but highly relies on an inquiry image. Therefore, it is reasonable to see why a low GtR is obtained. - Figure 5 demonstrates that different identities (such as different genders) can be obtained with different m, even with the same input query. There seems to be no way to control the "number of identities" generated from this model. If so, how was the supervised loss applied to train a face recognition model? - How can one ensure high inter-class and large intra-class variations as required for SFR? - B.3.3. The assertion that high-quality data is not indispensable for achieving markedly accurate facial recognition performance is somewhat counterintuitive and perplexing. - The method's reproducibility raises concerns, particularly with respect to the training of the model, which lacks clarity. Specifically, the functions F_1 and F_2 in Equations (6) and (7), as well as the role of C_att, are not explicitly defined, and these elements are absent from Figure 3. The proposed model generally lacks controllable factors to generate true synthetic face images that favor high inter-class and intra-class variations.

Questions

See above

Rating

5

Confidence

4

Soundness

2

Presentation

2

Contribution

3

Limitations

No issues are found in this aspect.

Reviewer L3p95/10 · confidence 4/52024-07-09

Summary

The paper proposes a novel approach named C to address privacy concerns in face recognition technology. The authors propose CemiFace, a diffusion-based method that generates synthetic face images with controlled similarity to a subject's identity center, enhancing the discriminative quality of the samples. This approach allows for the creation of diverse and effective datasets for training face recognition models without the need for large-scale real face images, thus mitigating privacy risks. CemiFace outperforms existing synthetic face recognition methods, significantly reducing the performance gap compared to models trained on real datasets. The paper also discusses the potential limitations and privacy implications of the approach, highlighting the need for ethical considerations in synthetic face generation for face recognition applications.

Strengths

1. The use of a diffusion-based model for generating semi-hard samples is an innovative approach that has not been extensively explored in the field of face recognition. 2. The approach can be extended to use unlabeled data for training, which is an advantage over previous methods that often require some form of supervision.

Weaknesses

1. The paper is not well organized. This paper should be reorganized to make it easier for the reader to understand the contributions and technical details of this paper. 2. Eq. 10 seems to be inconsistant to its description. According to the description, it is highly related to the time step. 3. Fig. 3 is hard to understand. The training losses are not illustrated in the figure. 4. Despite aiming to reduce privacy issues, CemiFace still uses a pre-trained model that could have been derived from datasets without user consent, raising ethical and privacy concerns.

Questions

Refer to weakness

Rating

5

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

Refer to weakness

Reviewer fJPL5/10 · confidence 4/52024-07-12

Summary

The paper titled "CemiFace: Center-based Semi-hard Synthetic Face Generation for Face Recognition" addresses a critical issue in face recognition (FR) related to privacy and performance degradation when using synthetic face images. The authors propose a diffusion-based approach, CemiFace, which generates facial samples with varying levels of similarity to an identity center. This method aims to enhance the discriminative quality of synthetic samples, thereby improving the performance of FR models trained on these datasets.

Strengths

Introducing Similarity controlling factor in synthetic face generation using a diffusion based approach.

Weaknesses

(a) Due to the introduction of this similarity control conditioning in the diffusion process there must be a change in total sampling time ( certainly it will also depend on the number of time steps considered in the diffusion process also) – An illustration/analysis on computational complexity of the proposed algorithm is needed. (b) Seems like the overall process is dependent on how(using which method) the value of m was determined during the diffusion process! (c) A complete pseudo-code on the proposed method would have helped the reader to understand the whole process. (d) Figure 3 could have been much more elaborated and in more details.

Questions

While comparing the proposed work with other SOTA methods - how did you generate the results of the SOTA method?

Rating

5

Confidence

4

Soundness

3

Presentation

3

Contribution

2

Limitations

Limitations were mentioned only in the last few sentences of the conclusion section but not otherwise stated separately.

Reviewer tciQ5/10 · confidence 4/52024-07-18

Summary

The paper proposes a new Face Recognition diffusion-based generation method. The diffusion process is completed with a semi-hard constraint on the synthetic reconstructed image: for each inquiry image of the (real) training set, the reconstructed image after the forward-backward diffusion process must have a specific cosine similarity with the inquiry image. As it is usual for such methods in Face Recognition, the resulting synthetic dataset is then used for training a Face Recognition model. This model is evaluated across diverse real datasets.

Strengths

The tackled problem is quite hard and needed at the same time. Current SOTA Face Recognition generation methods lead to a significant gap in terms of performance, compared to real Face Recognition datasets (of the same size). The idea of controlling the similarity to design semi-hard samples is also interesting.

Weaknesses

1) In Fig. 1, is the displayed similarity really the cosine similarity ? In the generated samples, the line with perfect similarity (equal to 1) seems to provide synthetic images which would not have a perfect similarity with the inquiry images displayed above the hypersphere. 2) [minor] In Eq. 3, the probability distribution of epsilon is not specified. 3) The authors should cite explicitly the works that use the training loss (Eq. 2) in this precise form, as there are alternative loss functions for diffusion models. A discussion on the reasons of this particular choice of diffusion loss might be a plus (e.g. in the appendix). 4) [minor] Although the lines 114-116 are accurate, they are misleading the reader. The widely known representation of Face Recognition embeddings is that they lie onto a hypersphere of dimension N, where each embedding is a point of the hypersphere. Those embeddings are clustered by identity on this sphere and the identity centers are roughly at the center of those clusters. The hypersphere mentionned in this paper is a hypersphere of dimension N-1, where the identity center is at the center of the sphere. 5) In lines 120-125, the authors should detail the range of similarities to the identity center, for each of the 5 splits of the CASIA training set. Only the average similarity of each split is specified. 6) [minor] Figure 3 should be a bit more explained than just its caption. 7) [major] Lines 155-162 are not well written and it is hard to understand how the margin m is used to guide the diffusion process. In particular, F_1 and F_2 are not defined, while some unused F is mentionned. C_sim seems to be a vector of unknown size. Also, the temporal guidance is too briefly described. 8) [minor] Some hyperparameters' values (alpha_t/beta_t, lambda) are not specified. 9) The right part of Fig. 4 displays two curves that do not have the same meaning for the x-axis. For AVG, the similarity is a constrained similarity (m) for training CemiFace (i.e. a similarity between a real inquiry image and a synthetic image). For CASIA, it is the similarity between one real image and its identity center (not a real image). To sum up, for AVG it is a similarity between 2 images, while for CASIA it is between 1 image and its identity center. Thus, comparing the two curves does not seem meaningful. 10) [major] The CosFace loss is used to train on synthetic datasets, while AdaFace is used to produce (identity-oriented) embeddings for the CemiFace training set generation. There should be only one model for both tasks, for fair comparisons. On Table 6, training on CASIA with AdaFace gives better results than with CosFace, so one could attribute the good performance of CemiFace to the fact that the authors used a stronger model (AdaFace) to generate the synthetic dataset than the model used to train on this dataset (CosFace). In addition, there should be a part studying the impact of this AdaFace choice (i.e. another loss), at least in the appendix. 11) The ROC curve on IJB-B/IJB-C for all synthetic methods of Table 6 would be a plus, as the accuracy is easily saturated, and not really used in industrial use-cases. Previous papers (related works) provide such ROC plots.

Questions

1) Could you explain the last sentence of Section 4.2.1 (lines 260-261) ? 2) In Section 4.2.2, why is the range of training m equal to [0,1] while the previous subsection concludes with an optimal range [-1,1] ?

Rating

5

Confidence

4

Soundness

3

Presentation

2

Contribution

3

Limitations

1) [major] The SimMat loss seems to be an interesting idea to lead towards m-similarity to the inquiry image, during training. But the derivation of the MSE loss (Eq. 3) assumes that the reconstructed image should be the inquiry image, and not a new image having a m-similarity with the inquiry image. I may be wrong here but I think that the diffusion loss of Eq. 3 is mathematically valid if the forward diffusion process is symetric to the reverse diffusion process, which is not the case here. 2) [major] In Section 4.2.1, there should be a part studying the difference between the required m and the estimated m post-training. That means that, for any required m, it is easy to compute the estimated m, i.e. the similarity between the resulting reconstructed image and the inquiry image. This estimated m must be quite different than the required m because Figure 4 shows that the best m is m=0, meaning that there is a 90 degrees angle between the inquiry image and the synthetized image. If m was truly equal to 0, the performance of the model would be very poor. So, there must have a difference between the required m and the real estimated m (post-training). This is also reflected in the conclusion saying that the best setting for m is to train CemiFace with m randomly sampled from [-1,1]. A similarity truly equal to -1 would bring a model with an astonishingly poor performance. ** Update ** I have increased my score from 4 to 5.

Reviewer fJPL2024-08-12

Reply to the Authors

Dear Authors, Thanks for your detailed explanation on all my queries. I don't have any further questions. With Regards, Reviewer fjPL

Authorsrebuttal2024-08-12

Dear Reviewer fjPL, Thank you very much for your kind reply. We have noticed that the rating has not been changed. Please kindly let us know if there is any additional concern we can clarify. We sincerely appreciate your valuable suggestions and will try our best to meet your criterion. Best regards, The authors of Paper 11025

Reviewer Crof2024-08-12

Thank you for the rebuttal. My questions have been answered. Also, after reading other reviews and rebuttals, it seems that the authors have sufficiently answered the questions. I hold my position as it is, as I believe it is a paper with a meaningful discovery on training with synthetic dataset.

Authorsrebuttal2024-08-13

Dear Reviewer Crof, Thank you for your appreciation of our work and your efforts in reviewing this manuscript. We will try to include the insightful discussion you suggested in the revised version. Best regards, The Authors of Paper 11025

Reviewer 5xKY2024-08-13

Thank you for your response. The authors have addressed my questions satisfactorily.

Authorsrebuttal2024-08-13

Dear Reviewer 5xKY, Thank you for your positive feedback in reviewing our paper. We will try to change the manuscript according to your suggestion in the updated version. Best Regards, The Authors of Paper 11025

Authorsrebuttal2024-08-13

Dear Reviewer L3p9, Thank you for your positive feedback in reviewing our paper. We will re-organize Fig.3 and the overall structure to make it easier for the reader according to your suggestions. Best Regards, The Authors of Paper 11025

Authorsrebuttal2024-08-13

Dear Reviewer tciQ, We are deeply grateful for the time and effort you spent reviewing our paper. We have carefully tried to address each of your questions based on your valuable feedback. Could you please take a brief moment (approximately 2-3 minutes) to review our responses? We appreciate your feedback, regardless of whether our answers have addressed your primary concerns. We are willing to provide additional details if you need further explanation. Best regards, The authors of Paper 11025

Authorsrebuttal2024-08-14

Dear Reviewer tciQ, We are deeply appreciated for your inspiring suggestion. We will try our best to improve the manuscripts in the updated version according to your opinion. Best Regards, The Authors of Paper 11025

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC