Dear reviewer,
Thank you for your interest and continual feedback. We appreciate the opportunity to further clarify the details of fMRI preprocessing pipeline. Regarding your questions:
Q1: Is reshaping into a 1D vector a common practice to use?
A1: Yes, indeed. A significant part of the neuroscience literature has worked with reshaped 1D fMRI data [1-5], including our baselines to be compared. By following these work to use a 1D representation, we ensure a direct, consistent and fair comparison.
Q2: Why do the reshaping but not use a 3D convolution?
A2: We definitely agree that 3D convolution is an interesting approach to explore in our future work. But in our current work, 1D convolution is also a reasonable solution. The reasons are as follows. The 3D structure of fMRI signals record not only activities of the cortical surface (grey matter), but also the subcortical areas and the white matter areas which are parts of the brain that are located below the cerebral cortex. However, for the visual reconstruction task, we are mainly concerned with the information from the cortical **surface**, which has been proven to be more relevant for conscious visual perception and representation in the brain [6, 7]. As a common operation taken in [1-5], the subcortical areas are removed during preprocessing which actually disrupts the 3D continuity of the data. With this preprocessing, 3D convolution would attempt to learn from disjointed spatial patterns and could potentially influence the meaningful signals from the surface regions. Moreover, 3D convolution increases the number of trainable parameters in the model than 1D convolution. With the limited data of the fMRI, possible overfitting is a concern. By using 1D convolution on the flattened data, we could limit model complexity while still retaining enough information to accurately decode the stimulus. We will diligently weave these detailed clarifications in the revised paper.
We deeply appreciate your insights and guidance. Should you have any further questions or require additional clarifications, please do not hesitate to propose them.
Warm regards,
Authors of Submission6209
[1] Z. Ren, J. Li, X. Xue, X. Li, F. Yang, Z. Jiao, and X. Gao, “Reconstructing seen image from
brain activity by visually-guided cognitive representation and adversarial learning,” NeuroImage,
vol. 228, 2021 (ref.4 in the paper)
[2] Y. Takagi and S. Nishimoto, “High-resolution image reconstruction with latent diffusion models
from human brain activity,” in CVPR 2023 (ref.5 in the paper)
[3] Z. Chen, J. Qing, T. Xiang, W. L. Yue, and J. H. Zhou, “Seeing beyond the brain: Masked
modeling conditioned diffusion model for human vision decoding,” in CVPR 2023 (ref.6 in the paper)
[4] M. Mozafari, L. Reddy, and R. van Rullen, “Reconstructing natural scenes from fmri patterns
using bigbigan,” IJCNN 2020 pp. 1–8. (ref.7 in the paper)
[5] F. Ozcelik and R. VanRullen, “Brain-diffuser: Natural scene reconstruction from fmri signals
using generative latent diffusion,” arXiv preprint arXiv:2303.05334, 2023. (ref.22 in the paper)
[6] Cecere, R., Bertini, C., & Ladavas, E. (2013). Differential contribution of cortical and subcortical visual pathways to the implicit processing of emotional faces: a tDCS study. Journal of Neuroscience, 33(15), 6469-6475.
[7] National Institute of Neurological Disorders and Stroke. Brain Basics: The Life and Death of a Neuron.(https://www.ninds.nih.gov/health-information/public-education/brain-basics/brain-basics-life-and-death-neuron) Accessed 3/19/2023.