Time Makes Space: Emergence of Place Fields in Networks Encoding Temporally Continuous Sensory Experiences

The vertebrate hippocampus is believed to use recurrent connectivity in area CA3 to support episodic memory recall from partial cues. This brain area also contains place cells, whose location-selective firing fields implement maps supporting spatial memory. Here we show that place cells emerge in networks trained to remember temporally continuous sensory episodes. We model CA3 as a recurrent autoencoder that recalls and reconstructs sensory experiences from noisy and partially occluded observations by agents traversing simulated arenas. The agents move in realistic trajectories modeled from rodents and environments are modeled as continuously varying, high-dimensional, sensory experience maps (spatially smoothed Gaussian random fields). Training our autoencoder to accurately pattern-complete and reconstruct sensory experiences with a constraint on total activity causes spatially localized firing fields, i.e., place cells, to emerge in the encoding layer. The emergent place fields reproduce key aspects of hippocampal phenomenology: a) remapping (maintenance of and reversion to distinct learned maps in different environments), implemented via repositioning of experience manifolds in the network’s hidden layer, b) orthogonality of spatial representations in different arenas, c) robust place field emergence in differently shaped rooms, with single units showing multiple place fields in large or complex spaces, and d) slow representational drift of place fields. We argue that these results arise because continuous traversal of space makes sensory experience temporally continuous. We make testable predictions: a) rapidly changing sensory context will disrupt place fields, b) place fields will form even if recurrent connections are blocked, but reversion to previously learned representations upon remapping will be abolished, c) the dimension of temporally smooth experience sets the dimensionality of place fields, including during virtual navigation of abstract spaces. Code for our experiments is available at1.

Paper

Similar papers

Peer review

Reviewer kdgt3/10 · confidence 4/52024-07-11

Summary

The paper shows that place cells can emerge in networks that autoencode temporally continuous sensory episodes based on spatially smoothed Gaussian random fields. The obtained place fields reproduce the disputed idea of remapping, the established fact that such spatial representations are uncorrelated, and a slow representational drift. The model implements “experience manifolds” in the network’s hidden layer and weakly spatially modulated (WSM) rate maps, which are interesting concepts that deserve more analysis. Also, dimensionality of the environment seems to be non-problematic, although also here a comparison to observations in biological experiments would be desirable.

Strengths

The paper introduces useful concepts such as experience manifolds or weakly spatially modulated (WSM) rate maps that may be useful in the further study of hippocampal function, although at the moment the ideas play a role only in the context of modeling, and the experience manifolds seem to be here merely metaphorical (although a similar analysis tool has been used in other instances of population codes). It is very good that clear prediction have been made, and that model is well describe (suppl. material).

Weaknesses

Main problems with the paper are the lack of strong quantitative evidence and the realizability of the model by biologically relativistic neural networks of a size comparable to CA3 (or much smaller considering that a mammal typically works with environmental information of much higher complexity. On the level of the current model, it could be discussed what model features are essential for what feature of the result. For more detail see Questions below.

Questions

Although it is possible to model CA3 as a recurrent autoencoder to show how place cells can emerge, can we really say that CA3 **is** an autoencoder or that it is its function to represent place fields? Would it be possible to reach the level where a quantitative comparison to experimental data becomes possible or is this difficult due to the ongoing discussion of what characteristics can be extracted from experimental data towards a meaningfully comparison? Can affirmative statements like “closely resembles results of experiments” (l295), “consistent with experimental results” (l203) be given more evidence? The numbers given (l232) are said to be “mirroring experiments with rodent CA3 place cells”, but isn't the absence of certain correlations both in the simulations and in the biological experiments rather weak evidence for the proposed model? Can any remainder correlations be compared? The fit in [47, Fig. 2A] is quite bold and does not represent the underlying mechanism, so that it is not exactly a good standard for comparison (l239). Would it nevertheless be useful to mention any quantitative agreement with even some interpretation of the biological data? The “experimental evidence that place cells can develop multiple place field” (l252) was never really “strong”. Would it be possible to consider more recent studies, so that a delicate analysis might be enabled to show whether the experimental observations are “mirrored” by the simulations? Can the bias in the literature be changed from classical papers towards more recent modeling approaches so that these are sufficiently discussed? It is not needed to do this here comprehensively, but the paper would gain from some comparison with other approaches as would the theory of spatial representations in mammals. It remains unclear why not the weakly spatially modulated (WSM) rate maps (or any other non-local representation of spatial information) can provide a similar autoencoding property and what specifically is necessary for the formation of roundish place fields. This question can be asked for most models that show place field emergence, so that here also theoretically not much progress is achieved, and it could seems that without the specifically designed noise of sigma=12cm (which is not varied here and not listed in the parameter tables in the appendix) the results may have been less realistic, while the diversity of realistic place fields in not achieved in the proposed model nor is it evaluated in comparison to biological data.

Rating

3

Confidence

4

Soundness

3

Presentation

3

Contribution

2

Limitations

The current version is limited in regards to the discussion of experimental finding, the effect of some of the model parameters, and of the realization of model by biologically more realistic neural networks. Other limitation and some open questions are addressed well in the paper.

Reviewer KNH26/10 · confidence 3/52024-07-12

Summary

The paper explores the emergence of place cells in neural networks by simulating the hippocampal area CA3, specifically when trained to recall and reconstruct temporally continuous sensory experiences encountered during navigation in simulated environments. The authors model this area as a recurrent autoencoder that operates on sensory inputs from simulated agents moving through environments with varying sensory landscapes. The results show place cells that resemble those recorded in the hippocampus.

Strengths

The approach is novel, and the idea of training a network to remember temporally continuous sensory episodes and then characterize its neural representations is a useful contribution. The paper conducts empirical evaluations in different types of environments (e.g., rooms with different shapes) and makes several testable predictions.

Weaknesses

The paper is written in a somewhat unusual style with the results following right after the introduction without a separate methods section making it difficult to follow the approach and understand the results. The interaction between sensory inputs and velocity integration seems to be missing.

Questions

How does your approach relate to path integration? Do you expect to see place cells in the absence of visual inputs? What exactly are the inputs to the place cells? Can those pixel-level visual inputs preprocessed with some sensory processing modules? Do you observe the remapping and changing shape of the place fields when the environment changes size or the walls move (e.g., O'Keefe & Burges 1996)?

Rating

6

Confidence

3

Soundness

3

Presentation

2

Contribution

3

Limitations

The authors should address the limitations more explicitly.

Reviewer 6VUu7/10 · confidence 4/52024-07-12

Summary

This paper presents a novel approach to understanding the emergence of place fields in the hippocampus. The authors propose that place cells can emerge from networks trained to remember temporally continuous sensory episodes, without explicit spatial input. They model the hippocampal CA3 region as a recurrent autoencoder (RAE) that reconstructs complete sensory experiences from partial, noisy observations. The model reproduces key aspects of hippocampal phenomenology, including remapping, orthogonality of spatial representations, and slow representational drift. The paper offers several testable predictions and provides a fresh perspective on the origin of place fields, suggesting that "time makes space" in neural representations.

Strengths

- Novel approach: The paper presents an intriguing hypothesis about the emergence of place fields from temporally continuous sensory experiences. - Comprehensive modelling: The model reproduces multiple key aspects of hippocampal phenomenology. - Testable predictions: The paper offers concrete predictions that could guide future experimental work. - Thorough experimentation: The authors conduct a wide range of simulations to test their hypotheses. - Biological plausibility: The model is grounded in known hippocampal anatomy and physiology. - Effective explanations: The paper employs (impressively) clear (to me) expressions and explanations that enhance its readability and impact.

Weaknesses

- Limited quantitative comparison with actual neural data from rodent studies. - Reliance on simplifying assumptions about sensory input structure (smooth Gaussian random fields). - Insufficient exploration of network parameter dependencies. - Absence of publicly available code for verification and extension of the work. I read the author’s justification, but this remains a weakness in my point of view. - Lack of formal theoretical analysis to explain the emergence of place fields, i.e. a rigorous mathematical framework that provides a deep analytical understanding of why place fields emerge in this model. - Inadequate positioning within existing frameworks of sequential data modelling, for example in relation to simple hidden Markov models (HMMs). - Although minor, the term "weakly spatially modulated" signals, which is important for understanding this work, lacks a clear definition. - Also minor, the section structure could be improved for clarity, particularly Section 2 which combines methods and results. How about calling it Experiments?

Questions

- The most important question for me is, how sensitive are your results to changes in network architecture and hyperparameters? - How does your model specifically relate to and extend beyond HMM in the context of sequential data modelling? - How about my 2 minor points in the weaknesses? - Have you considered conducting a more rigorous statistical comparison with experimental rodent data? If so, what challenges did you face? - Could you elaborate on how your model might be extended to account for other hippocampal phenomena, such as theta phase precession or grid cells?

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors have addressed some limitations of their work, particularly regarding the simplifying assumptions of their model. However, they could improve by: - Providing a more detailed discussion of the limitations of using smooth Gaussian random fields to model sensory inputs. - Addressing potential limitations in the generalizability of their findings to real neural systems. - Discussing any computational limitations or scalability issues of their approach. - Considering potential negative societal impacts, if any, of their work (e.g., implications for AI systems that might use similar principles).

Authorsrebuttal2024-08-13

thank you

Thank you for taking the time to read our paper and our responses. We are also grateful for the improved score. Regarding the point "Reliance on simplifying assumptions about sensory input structure", we intend to answer it together with our definition of WSM signals. Specifically, we verified the validity of using GRFs as models for WSMs in a parallel study but have not included it here to maintain the focus of this paper. We tested responses to visual stimuli at different spatial locations in VR-simulated rooms. We capture images at different locations in these rooms and pass them through several network models of visual systems. The features generated by these models consistently produce WSM fields. Thanks, Authors

Reviewer hM858/10 · confidence 5/52024-07-15

Summary

This study demonstrates that place cells can develop in networks trained to remember temporally continuous sensory episodes. The model CA3 as a recurrent autoencoder that recalls and reconstructs sensory experiences from noisy and partially occluded observations by agents traversing simulated arenas. The autoencoder training, which included a constraint on total activity, led to the emergence of place cells with spatially localized firing fields. These place cells exhibited key hippocampal characteristics: remapping, orthogonality of spatial representations, robust place field formation in variously shaped rooms, and slow representational drift. The authors present a unique framework of the optimal encoding of the experience space The study suggests that continuous spatial traversal results in temporally continuous sensory experiences, making several testable predictions about place field behavior under different conditions.

Strengths

This work presents a new perspective on the topic of place cell formation in CA3 during navigation. Their model is clear and well described model. The analysis of their network, combined with the predictions they make regarding hippocampal remapping, make this a relevant work for the field.

Weaknesses

More in-depth visualization in figure 3 would be nice, I like this framing of the problem in the text. Maybe show each example (suboptimal encoding, optimal encoding, remapping in a new environment, returning to the original environment) as a figure panel? It would be helpful for the authors to examine which of their assumptions and initializations are critical for the PF emergence they observe. What do your results look like when you use different history buffer, different levels of noise, a form of input different than the WSM, etc.?

Questions

I'm surprised that the units learn (via the recurrent weights) return to their initial positions - especially after changes (albeit small) in the input weights. My naive assumption would be that there are many redundant solutions which could accurately autoencode the experience vectors. Why does the network return to the same solution it initially had?

Rating

8

Confidence

5

Soundness

4

Presentation

3

Contribution

3

Limitations

I think this work would greatly benefit from greater comparison to other models of place field/cognitive map formation (in either an autoencoder or predictive learning framework), Levenstein et al. 2024 as a recent example.

Authorsrebuttal2024-08-13

Thank you so much for taking the time to read our paper, and for the positive evaluation.

Authorsrebuttal2024-08-10

Additional comments on "mirroring experiments with rodent CA3 place cells"

We are writing to provide more information about the comparison between our model and experiment, especially with regard to the correlation of the population within and between rooms during different visits. We propose to replace the original sentences with the following paragraph: We compared the correlation between rooms from cycle 2 and cycle 3, a scenario similar to the experiments in Alme et al [10]. The mean correlation between different rooms is $0.164 \pm 0.029$, and that of the same rooms is about 0.55 greater, $0.710 \pm 0.097$. The corresponding experimental values reported in [10] are also different by about 0.57: $0.08 \pm 0.005$ and $0.65 \pm 0.02$, respectively. Note that we should not expect precisely the same values of the correlation because the precise setups of the environments and experiments are different. For example, we have many trial rooms in our {\it in silico} study, while there are only 2 rooms in [10]. Furthermore, our network contains 1000 units while Alme et al. recorded only 342 neurons. Overall, we find that the population vectors of familiar rooms have significantly higher correlations as compared to different rooms, consistently with experiments.

Reviewer KNH22024-08-12

I appreciate the responses and clarifications and I adjusted my score accordingly.

Authorsrebuttal2024-08-13

We greatly appreciate your time spent reviewing our work and the increase of your score from 5 to 6. Thank you, Authors.

Reviewer 6VUu2024-08-13

I would like to thank the authors for their informative responses. There was a point that the authors listed but I think they forgot to address, namely: "Reliance on simplifying assumptions about sensory input structure." However, most of my other points, including my most major concern, were indeed addressed, so I adjusted my sore accordingly. Good luck!

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC