Summary
Generating simulated scATAC-seq data is important for developing new methods and gaining a deeper understanding of the data. However, the simulation is challenging due to dropout and high noise in the data. Authors proposed a diffusion + VAE type of method to solve the problem. The general idea is first to use VAE to project original scATAC data to a lower embedding space. Then impose a diffusion process in the embedding space. The lower embedding space is a GMM rather than the classical isotropic Normal, which is the novelty part of the method. This configuration makes biological sense as cells can be grouped as different cell types. Due to the introduction of GMM distribution in the hidden GMM space, there exists complications in generalising the diffusion loss function. The authors have shown nice and solid derivations in the appendix. The method is then applied to three datasets on three different tasks and achieved comparable performance as SOTA methods. It seems that authors have provided a convincing solution to the research question. While this paper is clearly written and the general idea is relatively easy to follow, it will be nice if the author could help to answer the following questions.
1. In Eq. 11, what is the parametric form of q_\phi(z|x_0)
2. In Eq. 14, what is the actual meaning of conditional information y, could you list the parameters
3. Could you provide a concrete network architecture of your network in a supplementary figure, i.e. including the tensors and their dimensions?
4. In sec 4.4.1, what is the dropout rate distribution in your simulated data, are they similar to the real data?
Depending on the answers, I may change my ratings in the future.
Strengths
Generating simulated scATAC-seq data is important for developing new methods and gaining a deeper understanding of the data. However, the simulation is challenging due to dropout and high noise in the data. Authors proposed a diffusion + VAE type of method to solve the problem. The general idea is first to use VAE to project original scATAC data to a lower embedding space. Then impose a diffusion process in the embedding space. The lower embedding space is a GMM rather than the classical isotropic Normal, which is the novelty part of the method. This configuration makes biological sense as cells can be grouped as different cell types. Due to the introduction of GMM distribution in the hidden GMM space, there exists complications in generalising the diffusion loss function. The authors have shown nice and solid derivations in the appendix. The method is then applied to three datasets on three different tasks and achieved comparable performance as SOTA methods. It seems that authors have provided a convincing solution to the research question. While this paper is clearly written and the general idea is relatively easy to follow, it will be nice if the author could help to answer the following questions.
1. In Eq. 11, what is the parametric form of q_\phi(z|x_0)
2. In Eq. 14, what is the actual meaning of conditional information y, could you list the parameters
3. Could you provide a concrete network architecture of your network in a supplementary figure, i.e. including the tensors and their dimensions?
4. In sec 4.4.1, what is the dropout rate distribution in your simulated data, are they similar to the real data?
Depending on the answers, I may change my ratings in the future.
Weaknesses
Generating simulated scATAC-seq data is important for developing new methods and gaining a deeper understanding of the data. However, the simulation is challenging due to dropout and high noise in the data. Authors proposed a diffusion + VAE type of method to solve the problem. The general idea is first to use VAE to project original scATAC data to a lower embedding space. Then impose a diffusion process in the embedding space. The lower embedding space is a GMM rather than the classical isotropic Normal, which is the novelty part of the method. This configuration makes biological sense as cells can be grouped as different cell types. Due to the introduction of GMM distribution in the hidden GMM space, there exists complications in generalising the diffusion loss function. The authors have shown nice and solid derivations in the appendix. The method is then applied to three datasets on three different tasks and achieved comparable performance as SOTA methods. It seems that authors have provided a convincing solution to the research question. While this paper is clearly written and the general idea is relatively easy to follow, it will be nice if the author could help to answer the following questions.
1. In Eq. 11, what is the parametric form of q_\phi(z|x_0)
2. In Eq. 14, what is the actual meaning of conditional information y, could you list the parameters
3. Could you provide a concrete network architecture of your network in a supplementary figure, i.e. including the tensors and their dimensions?
4. In sec 4.4.1, what is the dropout rate distribution in your simulated data, are they similar to the real data?
Depending on the answers, I may change my ratings in the future.
Questions
Generating simulated scATAC-seq data is important for developing new methods and gaining a deeper understanding of the data. However, the simulation is challenging due to dropout and high noise in the data. Authors proposed a diffusion + VAE type of method to solve the problem. The general idea is first to use VAE to project original scATAC data to a lower embedding space. Then impose a diffusion process in the embedding space. The lower embedding space is a GMM rather than the classical isotropic Normal, which is the novelty part of the method. This configuration makes biological sense as cells can be grouped as different cell types. Due to the introduction of GMM distribution in the hidden GMM space, there exists complications in generalising the diffusion loss function. The authors have shown nice and solid derivations in the appendix. The method is then applied to three datasets on three different tasks and achieved comparable performance as SOTA methods. It seems that authors have provided a convincing solution to the research question. While this paper is clearly written and the general idea is relatively easy to follow, it will be nice if the author could help to answer the following questions.
1. In Eq. 11, what is the parametric form of q_\phi(z|x_0)
2. In Eq. 14, what is the actual meaning of conditional information y, could you list the parameters
3. Could you provide a concrete network architecture of your network in a supplementary figure, i.e. including the tensors and their dimensions?
4. In sec 4.4.1, what is the dropout rate distribution in your simulated data, are they similar to the real data?
Depending on the answers, I may change my ratings in the future.
Limitations
Generating simulated scATAC-seq data is important for developing new methods and gaining a deeper understanding of the data. However, the simulation is challenging due to dropout and high noise in the data. Authors proposed a diffusion + VAE type of method to solve the problem. The general idea is first to use VAE to project original scATAC data to a lower embedding space. Then impose a diffusion process in the embedding space. The lower embedding space is a GMM rather than the classical isotropic Normal, which is the novelty part of the method. This configuration makes biological sense as cells can be grouped as different cell types. Due to the introduction of GMM distribution in the hidden GMM space, there exists complications in generalising the diffusion loss function. The authors have shown nice and solid derivations in the appendix. The method is then applied to three datasets on three different tasks and achieved comparable performance as SOTA methods. It seems that authors have provided a convincing solution to the research question. While this paper is clearly written and the general idea is relatively easy to follow, it will be nice if the author could help to answer the following questions.
1. In Eq. 11, what is the parametric form of q_\phi(z|x_0)
2. In Eq. 14, what is the actual meaning of conditional information y, could you list the parameters
3. Could you provide a concrete network architecture of your network in a supplementary figure, i.e. including the tensors and their dimensions?
4. In sec 4.4.1, what is the dropout rate distribution in your simulated data, are they similar to the real data?
Depending on the answers, I may change my ratings in the future.