Extraction and Recovery of Spatio-Temporal Structure in Latent Dynamics Alignment with Diffusion Models

In the field of behavior-related brain computation, it is necessary to align raw neural signals against the drastic domain shift among them. A foundational framework within neuroscience research posits that trial-based neural population activities rely on low-dimensional latent dynamics, thus focusing on the latter greatly facilitates the alignment procedure. Despite this field's progress, existing methods ignore the intrinsic spatio-temporal structure during the alignment phase. Hence, their solutions usually lead to poor quality in latent dynamics structures and overall performance. To tackle this problem, we propose an alignment method ERDiff, which leverages the expressivity of the diffusion model to preserve the spatio-temporal structure of latent dynamics. Specifically, the latent dynamics structures of the source domain are first extracted by a diffusion model. Then, under the guidance of this diffusion model, such structures are well-recovered through a maximum likelihood alignment procedure in the target domain. We first demonstrate the effectiveness of our proposed method on a synthetic dataset. Then, when applied to neural recordings from the non-human primate motor cortex, under both cross-day and inter-subject settings, our method consistently manifests its capability of preserving the spatiotemporal structure of latent dynamics and outperforms existing approaches in alignment goodness-of-fit and neural decoding performance.

Paper

Similar papers

Peer review

Reviewer 1oNu7/10 · confidence 3/52023-06-30

Summary

The authors address the problem of aligning behavior-related neural population dynamics, either within-subject but across different experimental sessions, or between subjects. This problem exists due to inter-subject variability in terms of which neurons are recorded or drift in the recording and is an important problem for systems neuroscience and for the development of brain-machine interfaces. Importantly, neural activity dynamics in many regions of the brain during behavior lie on a low-dimensional manifold and therefore have a defined spatio-temporal structure. While many state-of-the-art alignment approaches do not take into account this structure and therefore do not preserve it, the authors introduce a diffusion-guided method that does. The approach first uses a diffusional model to discover the manifold on which neural activity evolves (ie the spatiotemporal structure), then it uses this model to guide the alignment, which is done using MLE.

Strengths

Originality + quality: The authors developed a novel approach for time-series alignment that also discovers latent structure (the low-D manifold on which the neural activity evolves). Quality: The authors validate their model against a number of state-of-the-art alignment methods on both synthetic and real-world data, demonstrating the practical applicability of their method. Clarity: The authors clearly state the problem and its details, as well as how their method differs from existing ones, and its advantages. Significance: The method has significance to the alignment of timeseries with latent spatiotemporal structure, which is broadly relevant in systems neuroscience. Although the authors mainly focus on behaviorally relevant neural data, timeseries in other fields also often possess lower-dimensional latent structure, so this method could be broadly applicable. The authors approach could also be relevant to identifying latent structure outside of alignment context, although I am not familiar with how it compares to existing approaches to do so.

Weaknesses

Clarity: Figure and table legends should be more clear (see Questions section) Originality: Authors should state whether existing methods for discovering latent structures using diffusional models exist, and how their methods (e.g. architecture) differs from existing methods. Quality: The method is validated in data with strong, low-dimensional latent structure (monkey reaching tasks). The generalizability to other types of neural dynamics time series of varying dimensionality should be evaluated to determine limitations.

Questions

Figure 3B: Can the authors describe better what the takeaway from this panel is? In what ways does ERDiff improve on JSDM. The legend should tell the reader what the pink dots are (presumably critical points?). Figure 4: Please note in the legend or figure that the R-squared value is in %. I was confused at first as to how the R-squared could be negative before reading Table 1. Table 1: It is unclear to me what the denominator for the R-squared % calculation is. Can the authors make it more apparent in the table legend? If the time-series do not belong on a well-defined low-dimensional manifold, will the method hallucinate something? How well does the method work for more unstructured data? As the dimensionality of the latent structure increases, how much more badly does the method perform? The authors should quantify this.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

3 good

Contribution

3 good

Limitations

The authors should address in more detail the kinds of latent structures/time series that would pose more of an issue for their method. Otherwise, limitations are adequately addressed.

Reviewer xrSS6/10 · confidence 3/52023-07-04

Summary

Brain-computer interfaces require recalibration to accommodate drifts in the recorded neural populations over time. While there has been some success in aligning neural recordings based on their latent dynamics, deep learning-based alignment methods have also gained popularity since they ignore some of the implicit assumptions made by latent space methods, providing additional flexibility. However, deep learning neural alignment often ignores the temporal structure of the dynamics. The proposed model overcomes these limitations by first extracting the spatio-temporal structure in the source domain via a diffusion model and then aligning the target domain to the source dynamics. The authors demonstrated the success of the method in simulated and neural data, where it outperforms alternative alignment methods.

Strengths

The paper is adequately written and technically sound. The method was tested and shown to work well when applied to a simulated dataset and neural data. The method introduced here uses deep learning-based alignment while still exploiting the temporal dynamics that are critical in neural datasets. Robust alignment of neural recordings is crucial for the success of BCI applications, and in this work, they showed how this method outperforms existing alignment methods in both synthetic and neural datasets. Moreover, they also showed that the alignment can be performed not only across sessions of the same animal but also across animals.

Weaknesses

While the authors show the promise of their method to align neural responses, they overlooked other methods based on the alignment of latent dynamics that have been shown to be successful for BCI applications, as mentioned in the introduction. I believe that a systematic comparison to such methods is critical to fully grasp the significance of this work. For example, CCA, multiset CCA, hyperalignment, or Procrustes-based alignment. In the context of latent space alignment methods, another relevant piece of literature is the method introduced in (https://proceedings.neurips.cc/paper/2021/hash/aad64398a969ec3186800d412fa7ab31-Abstract.html), which also uses neural dynamics for alignment. Deep learning-based methods allow for more expressive functions, but they often come with additional computational costs, data demands, and the need for careful hyperparameter selection. None of these limitations are addressed in the manuscript, nor is there an explicit comparison across methods (latent space vs. deep learning-based), which could help demonstrate the promise of the method for practical BCI applications. The authors minimally showed the effect of dataset size on performance, but the lowest dataset size tested still had dozens of trials, which could be unrealistic in most practical settings. Additionally, it would be important to report the training times as a function of dataset size, as long training times would render the approach useless for real-time alignment. The method defines the alignment between a single source and target dataset, but ideally, one would pool data across all sessions for BCI decoding. It would be interesting to note if the proposed method also allows for multi-session alignment. The authors showed the success of the approach in the context of a single data simulation. However, to fully assess the robustness of the method, they could further test it under different conditions, such as measurement or latent noise, dimensionality, or tasks.

Questions

It is unclear from the text how the training and testing are performed. In the sentence "During testing, we align the test neural data to the training neural data so that we can directly apply the velocity decoder," it is not clear whether the authors include test data for alignment. Additionally, the authors mentioned that behavioral data is used for alignment, but it is also used to evaluate the success of the approach via decoding. This raises the question of whether there is a circular evaluation of the method. References 9 and 10 cite the preprint and peer-reviewed versions of the same article.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

2 fair

Limitations

The authors should include a section or provide a clear description of the limitations and assumptions of the method. They should also address the computational cost, data demands, and potential implications of the presented work.

Reviewer 2uTS7/10 · confidence 4/52023-07-07

Summary

* One of the key challenges in analyzing neural recordings is the scalability of models that link behavior and neural population activity across recording sessions or in inter-subject settings. • When analyzing single-trial neural population activity, past studies have pointed out that neural activity can be understood in terms of low-dimensional latent dynamics. Such low-dimensional latent dynamics are helpful when visualizing neural profiles across different task conditions and trials. • Generally, existing methods try to align latent dynamics by minimizing the difference evaluated by the metric between source and target domains. The paper proposes a method to align the source and target domains of multivariate neural data by learning/capturing latent spatiotemporal structure in the source domain with a diffusion model and applying it as a prior on learning/capturing spatiotemporal structure in the target domain. • The authors applied their model to the non-human primate motor cortex, testing both cross-day and inter-subject recordings.

Strengths

* The authors motivate their approach clearly by arguing that naively aligning time series using domain adaptation is ineffective as multivariate neural data has low SNR. Thus, leveraging low-dimension representation is a viable option. • The paper seeks to achieve a form of domain adaptation by aligning the spatiotemporal structure of latent dynamics of the target to the source using a novel alignment method (ERDiff). • The model and derivations are presented clearly. • An exhaustive comparison is provided showing that the their model outperforms standard models on both motor cortex and synthetic datasets.

Weaknesses

* The authors must clarify why diffusion models are necessary. How about considerably simpler two-step approaches -- like extracting latents with GPFA (Yu et al, 2009) and aligning them with the proposed ML alignment? * Alternatively how about comparisons with alternative approaches of comparable complexity -- e.g. adversarial alignment with DANN (Ganin et al, 2015)?

Questions

* Overall: there are many approaches for extracting latent structure from time series data (GPFA, CEBRA, T-PHATE, CILDS) -- one could readily realign the latents extracted from these approaches with the second stage alignment algorithm. * Apriori, why would one expect the diffusion model to be more effective at aligning the latents than these other approaches? * Because data from non-human primates is used, please clarify whether appropriate IRB approvals were obtained. Or if this is only a secondary analysis of existing datasets, the approvals obtained in the original studies could be mentioned.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

* No significant negative societal impact envisioned -- the potential use in BCI may suggest a positive social impact.

Reviewer wx957/10 · confidence 2/52023-07-13

Summary

This paper proposes a distribution alignment method (ERDiff), which combines extraction of spatio-temporal structures in latent dynamics from the source distribution and maximum likelihood alignment procedure on the target domain. The proposed method was evaluated on both synthetic and real data (neural recordings from non-human primate motor cortex), outperforms other methods under both cross-day and inter-subject settings.

Strengths

- The proposed method is novel and is technically sold - Performance of the method is well demonstrated on real data under inter-session/subject setup suggest that it could be an important tool with potential broad use in many field not just in neuroscience.

Weaknesses

-I don’t have any major concerns. Although the methods has been shown to outperform some of the current techniques, advantage of the proposed approach is not well demonstrated. I would like to see some analysis on computational cost

Questions

- Although the authors clearly mentioned the limitation of alignment method based on pre-defined metric, it would be nice see how the proposed method performs compared to these.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

3 good

Contribution

3 good

Limitations

- Limitations is not addressed in the draft

Reviewer tfQ56/10 · confidence 2/52023-07-13

Summary

The paper focuses on aligning highly variable neural population activities across days and subjects to stabilize the learning proposes and advance applications such as the brain computer interface. The idea is to train a VAE on one dataset that is hopefully self-consistent and then using a VAE trained on a dataset from another day or subject align the conditional distributions (encoders) for the latent spaces maximizing the likelihood of the source domain latent space under the new encoder distribution. The difficulty is the need to model the marginal source distribution of the latent space, which the work does using a diffusion model. The approach is demonstrated in comparison with alternative models on a synthetic dataset and actual neural recordings.

Strengths

1. A well written paper (but the abstract). 2. An interesting application of diffusion models. 3. Potentially impactful in practice of the BCI, more work, including further evaluation, is needed here however.

Weaknesses

1. A niche application and demonstration. A cellular neuroscience focused paper with no additional effort to demonstrate a general applicability of the approach in evaluations. 2. Poorly written abstract, especially in contrast to the rest of the paper.

Questions

1. Is the code going to be released publicly? Looks like the success of the approach depends less on the high level probabilistic descriptions in the paper than on the details of the actual implementation. 2. Is the data going to be released publicly for reproducibility?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

2 fair

Limitations

Potentially, the applicability of this work may be limited only to the demonstrated use in intracranial multicellular recordings, and it may not contribute to advancements in other areas of Machine Learning (ML). No demonstration was provided to counter this potential limitation.

Reviewer 9MfT6/10 · confidence 3/52023-07-18

Summary

Inter-individual and inter-session variability significantly complicate direct comparison of neural recordings collected over time, degrading trained behavioral models. This can be cast as a more general distribution alignment problem, shared across unsupervised learning. To address neural distribution alignment, the authors introduce a novel "ERDiff" method, which co-trains a variational autoencoder and a diffusion model to extract latent spatiotemporal structure in a source dataset. To align with a desired target dataset, parameter finetuning is performed on the read-in layer of the probabilistic encoder learned during training to match the target dataset. Simulation and experimental results suggest that ERDiff captures relevant spatiotemporal structure, performing competitively against other baseline methods including those based on overall minimizing distribution divergence as well as based on adversarial methods. Overall, the authors argue that ERDiff is better able to capture the important spatiotemporal structure of their trialwise data; for example, from a monkey center-out reaching task in the experimental results. This focus differentiates ERDiff from many other alignment methods, which ignore unfolding dynamics in their alignments.

Strengths

Incorporating spatiotemporal structure into the alignment of neural recordings is a novel approach, as these methods traditionally consider successive data points as independent samples or learn a set of low-dimensional latent dynamics which can then be aligned. These dynamics are particularly critical where the the trialwise dynamics strongly influence both behavior and neural activity over time. By directly learning and aligning the latent spatiotemporal structure, ERDiff shows stable performance even over relatively low sampling density, retaining relatively high decoding performance compared to other baseline methods even with decreasing numbers of trials in the target domain. These properties suggest that it is also likely that the general ERDiff approach may be useful in other cases where distribution shift between a source and target domain obscures but does not remove a shared latent structure. The general approach of combining variational autoencoders (VAEs) and diffusion models (DMs) has been previously introduced (e.g., Panday et al., TMLR, 2023); however, ERDiff is a significantly different formulation of the idea and represents a novel approach in leveraging the relative strengths of these methods through cooperative training.

Weaknesses

While the current experiments extensively compare inter-session and inter-subject differences in real neural recordings --- in addition to the simulated experiments --- it is not clear to what extent the presented findings might generalize to data sources without such a clear trial structure. For example, in recordings collected during unconstrained exploration or sleep, low-dimensional latent factors may still drive a successful alignment. Nonetheless, it is not clear how ERDiff would best be deployed in that context. This is particularly relevant as the described real-data experiments additionally incorporated velocity information, and the performance of ERDiff without a behavioral signal during training is thus unclear. In the current paper, my additional concern with the experiments is on the relative baselines against which ERDiff is compared. It would be informative to see a direct comparison with canonical correlation analysis (CCA) approaches, which have been used to date in aligning neural datasets with a strong temporal correspondence (e.g., Gallego et al., Nat Neuro, 2020).

Questions

Would the authors be able to directly comment on the relationship between their work and other field-standard methods to identify low-dimensional, latent factors such as LFADS (Pandarinath et al., Nat Methods, 2018) ? In particular, the relative benefits of learning the low-dimensional latent structure as part of the alignment --- as compared to other existing methods which learn low-dimensional dynamics which can then be aligned --- is not clearly explained. This would help to better situate the work in the literature.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

Although I do not see a Broader Impact section included in the current submission, I do not anticipate potential negative societal impacts given the constrained focus of the work.

Reviewer 1oNu2023-08-10

Thank you for your detailed response addressing my concerns. Given the response as well as other reviewers comments, I maintain my opinion that the manuscript is suitable for publication and will keep my score at 7.

Authorsrebuttal2023-08-11

We appreciate the timely response. Thank you again for your evaluation and recognition of our work.

Reviewer tfQ52023-08-13

Thank you for your explanations and an improved abstract. I still hold that a generality can only be demonstrated in a wider set of experiments rather than hypothesized. I do value the potential uses that a method like the proposed can have if it works outside of the demonstrated domain, however as it stands, the evidence that it does is lacking. Nevertheless, this is a strong manuscript and an interesting approach that fits the "Neuroscience and cognitive science" section.

Authorsrebuttal2023-08-13

Additional Experiments on General Time-series Domain-Adaptation Datasets

We thank the reviewer for the valuable response. We truly agree that the evidence from experiments carries more weights in determining the generalizability of each method. Here we additional conduct experiments on two general time-series datasets widely used in domain-adaptation papers: (1) Boiler Fault Detection Dataset [1]. The dataset contains sensor data from three distinct boilers, collected between March 24, 2014, and November 30, 2016. Each boiler in this dataset is considered as a unique domain (we represent as 1,2, and 3). The objective of the learning task is to predict the faulty blowdown valve of each boiler. The results can be found in Table 1 below. (2) City Air Quality Forecast Dataset [2]. The dataset is composed of air quality, meteorological, and weather forecast data from three cities, denoted as A, B, and C. Each city is treated as a unique domain. Using both the air quality and meteorological data, our objective is to predict PM2.5 levels. The results can be found in Table 2 below. *Implementation Details:* Besides the three methods that focus on general time-series domain adaptation we compared in the manuscript, we add one more fundamental baseline: LSTM_S2T. This approach trains a vanilla LSTM model using source domain data and then directly applies it to the target domain without adaptation. This method represents the performance lower bound. In ERDiff, we apply $4$ STBlocks in the diffusion model (DM). For a fair comparison, we set the size of the latent dimension equal to the representation space size used in other methods. The batch size is set as 128. Owing to the strong domain learning capabilities of DM and our proposed corresponding maximum likelihood alignment (MLA) in adaptation phase, most times ERDiff outperforms existing methods in terms of alignment quality and it reaches the highest performance on average. | Method | 1$\rightarrow$2 | 1$\rightarrow$3 | 3$\rightarrow$1 | 3$\rightarrow$2 | 2$\rightarrow$1 | 2$\rightarrow$3 | Avg | | :--------: | :-------------: | :-------------: | :-------------: | :-------------: | :-------------: | :-------------: | :-------: | | LSTM_S2T | 67.09 | 94.54 | 93.14 | 56.09 | 84.99 | 91.31 | 81.19 | | SASA | 71.54 | 96.39 | **94.77** | 63.15 | 87.76 | 93.59 | 84.53 | | RDA-MMD | 73.95 | 96.30 | 94.14 | 65.05 | 88.11 | **94.42** | 85.34 | | DAF | 74.55 | **96.54*** | 94.58 | 65.03 | 88.85 | 94.19 | 85.59 | | **ERDiff** | **75.26*** | 96.13 | 94.14 | **66.66*** | **89.09*** | 94.02 | **86.21** | *Table 1: AUC Score ($\%$) on Boiler Fault Detection Dataset. $\star$ denotes significance p-value <0.02 compared with the best baseline.* | Method | B$\rightarrow$A | C$\rightarrow$A | A$\rightarrow$B | C$\rightarrow$B | B$\rightarrow$C | A$\rightarrow$C | Avg | | :--------: | :-------------: | :-------------: | :-------------: | :-------------: | :-------------: | :-------------: | :-------: | | LSTM_S2T | 40.20 | 48.91 | 52.81 | 68.14 | 13.82 | 13.82 | 39.62 | | SASA | 34.26 | 40.91 | 48.15 | 56.80 | 13.49 | 13.46 | 34.51 | | RDA-MMD | 32.98 | 37.88 | 45.42 | 52.78 | **13.19** | 13.18 | 32.57 | | DAF | 31.75 | 36.86 | 44.24 | 52.93 | 13.22 | 13.07 | 32.02 | | **ERDiff** | **31.05*** | **35.45*** | **43.30** | **51.36*** | 13.41 | **13.03** | **31.28** | *Table 2: RMSE on Cities Air Quality Forecast Dataset. $\star$ denotes significance p-value <0.02 compared with the best baseline.* We thank the reviewer for the kind comments. [1] used in: Time Series Domain Adaptation via Sparse Associative Structure Alignment. (Ruichu et al., 2021) [2] https://www.microsoft.com/en-us/research/project/urban-air/

Reviewer xrSS2023-08-15

I thank the authors for their really comprehensive response. I mostly agree with their comments and I have updated my score accordingly. I still believe that some of there comparisons, even if not tested, should be mentioned in the final version of the manuscript; which also emphasizes the relevance of this method, as they discussed here.

Authorsrebuttal2023-08-15

Thank you

We appreciate the reviewer for the kind response and constructive comments. As suggested, we will incorporate the systematic comparison across methods discussed here into the final manuscript.

Reviewer wx952023-08-17

I appreciate the authors for addressing my questions in detail. I will keep my original score at 7.

Reviewer 9MfT2023-08-17

I thank the authors for their comprehensive response in clarifying the contribution of this work, and I’ve updated my score correspondingly. In particular, I find the attentional experiments on (1) a rat hippocampus dataset and (2) when removing the behavioral signals from ERDiff particularly compelling. These broadly reinforce the author’s point that a trialwise structure—though not behavioral information—is necessary for a successful application of ERDiff. I appreciate the direct comparison with unsupervised-CCA, but I am still uncertain that this is the right baseline. Would it not be more meaningful to compare unsupervised-CCA with the ERDiff model without behavior signals of source domain during VAE training?

Authorsrebuttal2023-08-18

Thank you

We sincerely appreciate the reviewer's positive evaluation of our additional experiments and contribution. We apologize for the ambiguity and would like to clarify that the term 'unsupervised' in unsupervised-CCA denotes unsupervised neural distribution alignment. This refers to the exclusion of supervised behavioral labels (e.g., direction and velocity) in the **target domain** during the alignment phase. In the above experiments of unsupervised-CCA (U-CCA) we presented, the behavioral signals from the **source domain** are incorporated during VAE training. Hence, we compare it with the version of ERDiff that also incorporates behavior signals of source domain. We thank the reviewer once more for the valuable response and suggestions.

Reviewer 2uTS2023-08-20

I thank the authors for the detailed clarifications and additional experiments, and have updated my score accordingly. I have no further questions for the authors.

Program Chairsdecision2023-09-21

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC