Reasons to reject
- The scope of the paper is not clear from the Title, Abstract and Introduction. The paper presents itself as a survey of synthetic data in general, but its scope is more narrow: synthetic text data mainly generated using hand-engineered templates and algorithms (but not generative algorithms). This is not a bad thing : a review of synthetic text data can still be very useful. However, clarifying the scope is essential to reaching the intended audience, because what most people understand by « synthetic data » is artificial data generated by an AI model, e.g., a GAN or a Bayesian network, which was trained on real data.
- The paper is very extensive on the applications of synthetic text data, but is missing a taxonomy/categorization of the synthetic data generation models. The introduction does not delineate the boundary between human-aided techniques (such as algorithms/simulations) and fully data-driven synthetic data (« rather than being directly created by human. » ). A table of available synthetic datasets grouped by method and/or use case is missing.
- The methodology used to identify the tasks to which synthetic text data is applied is not clear. It is not clear if the review is exhaustive of the field, and how « good » use cases are delineated from « bad » use cases.
Questions to authors
Questions:
- How does your work relate to other surveys in this space, e.g., https://arxiv.org/abs/2205.03257?
Main comments:
- Please consider adding formal definitions and corresponding terms for the various types of synthetic data considered in this work, such as template-generated, simulated, algorithm-generated synthetic data, etc.
- The statement that synthetic data can create « anonymized » datasets « that do not contain sensitive personal information » is incorrect. An important missing reference is Stadler et al. [A] which demonstrates that synthetic data is not immune to inference attacks (such as membership inference and attribute inference), one criteria for anonymisation as per the Working Party 2014’s Opinion on Anonymization techniques. The notion that synthetic data is inherently privacy-preserving has been seriously challenged (see also Houssiau et al, 2022 [B] and Meeus et al., 2023 [C]). Unless formal privacy protections are put in place (differential privacy guarantees), synthetic data will leak private information, especially about vulnerable records. How to achieve synthetic data that is truly privacy-preserving while remaining useful is an open problem.
- How much of the literature in this space relies on template/algorithm-generated synthetic data, with prior knowledge infused by humans and how much of it uses LLM-generated synthetic data ? When is it better to use one over the other?
- Please consider adding a table with the datasets which are publicly available. This could help this paper become a useful starting point for identifying already existing synthetic text datasets.
- The paper surveys mainly results from the last couple of years, however synthetic data similar to the focus of this work has already been used in the past, see e.g., Kocijan et al.[D]’s dataset for solving the Winograd Schema Challenge. Please consider including citations from older work and incorporating their contribution in the review.
- The question asked in the «emergent self-improvement capability » paragraph seems to have already been (negatively) answered in Shumailov et al. [E] Please consider including a discussion of this paper and updating the conclusions in light of their finding.
[A] Stadler, T., Oprisanu, B., & Troncoso, C. (2022). Synthetic data–anonymisation groundhog day. In 31st USENIX Security Symposium (USENIX Security 22) (pp. 1451-1468).
[B] Houssiau, F., Jordon, J., Cohen, S. N., Daniel, O., Elliott, A., Geddes, J., ... & Szpruch, L. (2022, October). TAPAS: a Toolbox for Adversarial Privacy Auditing of Synthetic Data. In NeurIPS 2022 Workshop on Synthetic Data for Empowering ML Research.
[C] Meeus, M., Guepin, F., Creţu, A. M., & de Montjoye, Y. A. (2023, September). Achilles’ heels: vulnerable record identification in synthetic data publishing. In European Symposium on Research in Computer Security (pp. 380-399). Cham: Springer Nature Switzerland.
[D] Kocijan, V., Cretu, A. M., Camburu, O. M., Yordanov, Y., & Lukasiewicz, T. (2019, July). A Surprisingly Robust Trick for the Winograd Schema Challenge. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 4837-4842).
[E] Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., & Anderson, R. (2023). The curse of recursion: Training on generated data makes models forget. arXiv preprint arXiv:2305.17493.