Summary
This paper introduces Brain Diffusion for Visual Exploration (BrainDiVE) that aimed at exploring the fine-grained functional organization of the human visual cortex. Motivated by the limitations of previous studies that relied on researcher-crafted stimuli, BrainDiVE leverages generative deep learning models trained on large-scale image datasets and brain activation data from fMRI recordings. The proposed method uses brain maps as guidance to synthesize diverse and realistic images, enabling data-driven exploration of semantic preferences across visual cortical regions. By applying BrainDiVE to category-selective voxels and individual ROIs, the authors demonstrate its ability to capture semantic selectivity and identify subtle differences in response properties within specific brain networks. It is also shown that BrainDiVE identifies novel functional subdivisions within existing ROIs, highlighting its potential for providing new insights into the human visual system's functional properties.
Strengths
The paper's methodology shows promise in capturing semantic selectivity and identifying fine-grained functional distinctions within visual cortical regions. Also, the paper's potential significance lies in applying BrainDiVE to understand the fine-grained functional organisation of the human visual cortex.
By providing insights into category selectivity, response properties, and sub-regional divisions, the paper opens avenues for further exploratory neuroscience studies.
To achieve this, the authors perform the experiments in the manuscript are extensive, covering:
1. the semantic specificity of the method by decoding images from task fMRI and literature-obtained ROIs
2. compared the abstraction of face representation in the brain by comparing images decoded from two regions, the fusiform face area (FFA) and the occipital face area (OFA).
3. Use the method to extend knowledge of brain function by finding subdivisions in known areas in the cases of food decoding.
The authors describe the experimental setting clearly at the beginning of section 4 for all cases.
The experiments show that the manuscript has a good balance between the combination of known methodological approaches and addressing interesting questions in neuroscience. Specifically, the authors perform a commendable effort in obtaining quantitative results using human evaluators (cf Tables 3 and 4).
Finally, the authors show through results evidence that their method is a window into understanding region-specificity hypotheses of brain regions which are knowingly involved in different aspects of visual processing.
Weaknesses
The paper could benefit from comparing its results with previous contributions [e.g. 9 and 10] to highlight its novelty and contributions in relation to existing methods.
The methodology part lacks clarity, particularly in explaining key components like the diffusion model architecture and brain-guided synthesis—further clarifying the role of the image-to-brain encoder in influencing the denoising process during inference.
The generalisability and reliability of results are hard to asses through a small dataset, specifically just 10 subjects.
Furthermore, the paper lacks clarity on how the sub-divisions of the visual cortex are being verified or validated. Whether the results are specific to the analysed regions or the algorithm is being too biased by the experimental condition is not clear through the experiments. For instance, what would happen if the authors try to decode an area not specific to visual processing? An example of this would be using subsections of the orbitofrontal gyrus or other brain areas not expected to perform well in reconstructing visual stimuli.
As a small point authors should review the presentation of images. Figure 1 shows mostly best-case scenarios which are then not as good in Figure 4 (for instance the case of images generated from face voxels). This might bias readers. Second, some details, such as figure 4 appearing before figure 3 make reading the paper confusing.
Questions
1. Validation of BrainDiVE's Effectiveness:
a. Can the authors provide more details on the validation process to ensure that BrainDiVE-generated images indeed effectively activate the targeted brain regions?
b. Could the authors consider conducting more extensive comparisons with other existing methods [e.g. 9 and 10], to demonstrate the advantages and uniqueness of BrainDiVE in eliciting specific brain activations?
c. Could the authors show the results on decoding a baseline region which is not involved in visual processing?
2. Clarity in Methodology: The methodology part could benefit from more clarity and detailed explanations of key components, such as the diffusion model architecture and the exact implementation of brain-guided synthesis. Improving the architecture in Figure-2 might help.
3. Statistical Analysis of Qualitative Evaluation: As the paper relies on qualitative evaluation with 10 subjects, could the authors mention this explicitly and what limitations are expected from this in the limitations section?
4. Ethical Considerations: Could the authors describe the demographics of the population, or acknowledge the lack of this information to inform of wether the results are biased to a specific gender or population group.
5. Reproducibility: Are the authors intending to release the code and not publicly available data such as the scores produced by the human raters upon acceptance of the paper? This is key to guarantee reproducibility
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Limitations
Authors have not addressed the negative societal impact of their work, but this can be fixed by adding specific text in the Discussion section. The authors don't mention the demographic characteristics of the small human subject database that they have used not mention any ethical concerns related to decoding images from brain activations. This is fixable so authors should address it.