Dis-inhibitory neuronal circuits can control the sign of synaptic plasticity

How neuronal circuits achieve credit assignment remains a central unsolved question in systems neuroscience. Various studies have suggested plausible solutions for back-propagating error signals through multi-layer networks. These purely functionally motivated models assume distinct neuronal compartments to represent local error signals that determine the sign of synaptic plasticity. However, this explicit error modulation is inconsistent with phenomenological plasticity models in which the sign depends primarily on postsynaptic activity. Here we show how a plausible microcircuit model and Hebbian learning rule derived within an adaptive control theory framework can resolve this discrepancy. Assuming errors are encoded in top-down dis-inhibitory synaptic afferents, we show that error-modulated learning emerges naturally at the circuit level when recurrent inhibition explicitly influences Hebbian plasticity. The same learning rule accounts for experimentally observed plasticity in the absence of inhibition and performs comparably to back-propagation of error (BP) on several non-linearly separable benchmarks. Our findings bridge the gap between functional and experimentally observed plasticity rules and make concrete predictions on inhibitory modulation of excitatory plasticity.

Paper

Similar papers

Peer review

Reviewer TQSB6/10 · confidence 5/52023-06-14

Summary

The authors extend previous work on deep feedback control (DFC) learning to make the functional form of plasticity a more realistic reflection of what is observed biologically. In particular the authors: 1. Introduce a feedback control mechanism using targeted, neuron-specific inhibition signal that allows for synaptic plasticity that more closely resembles experimentally observed plasticity (e.g. BCM), compared to the error-based learning employed in a wide variety of previous papers. 2. Show that learning with the extended plasticity rule still comes relatively close to backpropagation-level performance on several benchmarks (MNIST, FMNIST). 3. Discuss at length the testable predictions of their extended model.

Strengths

The strengths of the paper are as follows: 1. Reducing an algorithm derived from optimization principles to the point that it can replicate LTD/LTP experiments is quite difficult, and many (but not all) preceding algorithms do not appear to succeed in this regard, relying more on error-based signals that are less clearly related to experimentally observed plasticity phenomena. 2. Even with biophysically motivated modifications, the algorithm still performs quite well on image classification tasks, which is also difficult. 3. The paper is very clearly written and organized.

Weaknesses

The authors themselves identify several key weaknesses, which I will elaborate on below. However, to me the principal weakness is that the contribution is very incremental: previous studies (e.g. Payeur et al. 2021) have demonstrated that related learning algorithms can, in certain regimes, resemble BCM-like learning, and the high performance and 'locality' of the DFC family of algorithms has already been explored extensively in previous studies (e.g. Meulemans et al. 2021 & 2022). Therefore, it seems to me that the key improvement this study demonstrates is that the DFC algorithms can, with some extra tweaks, also resemble this type of Hebbian learning. Other weaknesses: 1. As the authors note, their current learning scheme involves a controller with access to a highly complex Jacobian. This Jacobian is a function of the fixed-point of the network dynamics, and so as far as I can tell, for every individual stimulus, the network's feedback weights would have to be different in order to match the principled feedback control dynamics. Previous DFC studies demonstrated that it's possible to get away with simpler controllers with fixed weights, but this approximation is not used in this study, and it's not clear why--as is, though the learning rule locally resembles BCM, the feedback signal actually used to achieve high performance on MNIST/FMNIST is essentially almost as complicated than the backpropagation error signal itself. 2. The unrealistic 1-1 mapping between inhibitory and excitatory neurons makes it difficult to pin down where exactly the controller feedback should be expected to be at the level of a cortical microcircuit.

Questions

Could you elaborate on the relationship between this algorithm and the algorithm proposed in Payeur et. al 2021? In particular, that algorithm also replicates BCM-like LTD/LTP phenomena--it seems as though the ability to replicate BCM-like learning dynamics is not unique to the DFC family of models. What are the key differences in testable predictions between this model and Payeur 2021? Or older predictive coding-based models like Urbanczik & Senn 2014? Is the method employed in this study inherently unique to the DFC family of algorithms, or does it apply equally well to the other algorithms as well? E.g., could similar interneuron-modulated plasticity also be applied to predictive coding-based models? If the network is driven to a near-optimal performance regime by the controller for every stimulus, and this is occurring in a biological system, would there ever be any observable improvements in performance? Or would the animal be performing instantaneously very well, with the only observable progressive change being a reduction in the energy required for the controller? If this is true, which biological systems could this algorithm adequately model? In this model, is top-down feedback exclusively isolated to inhibitory neurons? Is this compatible or incompatible with models that propose feedback is also (or exclusively) directed to apical dendrites of pyramidal neurons?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

4 excellent

Presentation

4 excellent

Contribution

2 fair

Limitations

There are no obvious negative societal impacts of this work, and the authors very adequately address the limitations of their work.

Reviewer A5Rt7/10 · confidence 4/52023-06-25

Summary

The authors propose a neural plasticity mechanism with a key role for disinhibition. In particular, they apply a deep feedback control framework whereby feedback driven inhibitory neurons mediate changes in the feedforward excitatory connections. The proposed rule is argued to hold desirable properties in that it is local, captures/predicts experimentally observed plasticity, and can guide error-modulated learning.

Strengths

- the paper is generally well written (though there are some typos, see below) - the proposed plasticity is novel and well theoretically groundeded - the experimental predictions are well presented and appear feasible to perform

Weaknesses

- I am not fully convinced of the extent of novelty within this work. It seems to me that the authors made only slight alterations to the setup in [1, 2] such that inhibitory neurons are now included in the architecture (instead of + Qc we now have - (- Qc)) . Of course the brain does include inhibitory neurons, so adding them explicitly is arguably one step closer to plausibility, but I think the authors could present a stronger argument by addressing more the functional/computational differences/benefits with this addition compared to the previous models [1,2]. Comparing it to theorem 2 in [1], is the key difference that the interneurons enable the network to avoid the need to store/discriminate between the feedforward (ff) activity and the total activity of the excitatory neurons? If so I think this could be more clearly conveyed - I think the authors could more in relating their work to other computational works which consider the role of inibitory neurons on plasticity. For example, would [3] make different predictions in Fig 2? Moreover, the authors do not relate their model to [4], which to my knowledge seems to have significant intersection in terms of modeling and possibly predictions. - though I found the writing in general good, there were still a fair few sloppy errors and places which were unclear to me (see below) References: [1] Meulemans et al. 2021, Credit assignment in neural networks through deep feedback control [2] Meulemans et al. 2022, Minimizing Control for Credit Assignment with Strong Feedback [3] Sacramento et al. 2018, Dendritic cortical microcircuits approximate the backpropagation algorithm [4] Greedy et al. 2022, Single-phase deep learning in cortico-cortical networks

Questions

- Fig 1a: what's the dfiference between error and feedback signals? - Fig 1b: what is f and g? Why are they different? - Fig 1 caption: eqprop not defined - line 93: "While these error-modulated learning rules prove functionally useful...they fall short of capturing established properties of experimentally observed plasticity, such as a postsynaptic activity threshold". Forgive me if I'm being naive: what is meant here postsynaptic activity threshold? It's not clear to me. - line 99: "However, the model (burstprop) assumes a rigid circuit architecture to decode errors from multiplexed spike trains and thus does not generalize to other neuronal circuits". I don't understand the logic othis sentence: could the authors elaborate? - line 102: "..necessitate feedback signals to be weak.."; this sentence confused me a bit. It seems to me there's a difference between 1. the strength of feedback signals and 2. whether feedback signals activity influence neuronal activity. For example, in burstprop one might have a very high apical potential (and burst rate) but little/no influence on the event rate - line 130: define L - I found the explanation of how optimal feedback weights are found confusing (equation 7). Firstly, J contains L vectors for each u_i, but this is multiplied by only one u_i? Secondly, it seems equation 7 just shows how the change in r relates to any changes in u - this is true in whatever weights are chosen, I don't understand how this is an equation to be 'solved'. Finally, from what I understand in [1] it's true that that if the column space of Q is equal to the row space of J (line 145) then equivalence to Gauss-Newton optimisation is possible, but is this choice of Q necessarily optimal in this case? - equation 8: use of subindices meaning datapoints here whilst layers elsewhere is confusing - lines 152-159: it was only clear to me after looking at [2] that minimising the surrogate loss should also minimise the task loss. This should be presented more clearly - line 156: I don't see the logic of 'as a result' from the previous sentence; doesn't it just follow from equation 5? - line 167, 169: Eq. not equation - Fig 2: is the x axis post-synaptic rates? this should be in caption. Same with frequency in 2d - Fig 2b: inset diagram is unclear; is the black line the true inverse function? I am also curious to see an imperfect approximation where the approximation is above/below the true values at low/high values respectively - Fig 2 caption: 'resembles experimentally observed plasticity in simulated in vitro condition' - this seems a strong statement when only one paradigm (an isolated neuron) is actually shown - line 197: "The dis-inhibitory feedback signal creates a stable equilibrium state for the weight dynamics that coincides with the postsynaptic target computed by the feedback controller". I'm sorry, I didn't understand this sentence - for the student-teacher task in section 5, what is f? is it the output of a randomly intialised teacher network of the same architecture as the student network but without feedback? If without feedback, why does the solution for H in 3c bottom not go to zero? Also, is it necessary to set the feedback weights as the transposed Jacobian. Would it also work if its column space was equal to the row space of J? - For simulations in image tasks, n=3 seeds seems quite low to me. Could you repeat for a higher number? (like n=10) - For readability I'd recommend rotating table 2 (perhaps switching rows with columns is necessary) typos: line 105. Full stop after references. equation 8: bold Q line 203: appendix section 'xyz' (also correct in appendix itself) line 207: poor grammar line 244: appendix C and D Tabel 1 caption: full stop at end line 284: comma after 'fashion-MNIST' line 336: has -> have References: see refs for weaknesses above

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Limitations were well addressed by authors

Reviewer EPnb7/10 · confidence 3/52023-06-30

Summary

This paper uses adaptive control theory to derive plasticity rules for a fairly plasubile model of multi-layer networks in the brain, which are capable of matching the performance of backpropagation without restrictive assumptions such as mirrored or very weak feedback connections. Specifically, each excitatory neuron has a dedicated inhibitory neuron, which is in turn inhibited by the error signal propagated along feedback connections. The plasticity rule for excitatory units is essentially Hebbian but modulated by the inhibitory signal, matching in vitro plasticity experiments. A brief theoretical derivation of the rule is augmented with experimental results on both toy and relevant problems.

Strengths

* This work presents a relatively well-founded model for feedback of error in the brain, which seems to me one of the most biologically-plausible models of backprop to date. * All of its presentation is clear and should be accessible to readers with a wide variety of related backgrounds.

Weaknesses

* There are some lingering issues of biological plausibility: There are still significant constraints placed on the feedback weights; each excitatory neuron is assumed to have its own inhibitory neuron (which doesn't match the ratios found in the brain); and Dale's law is not enforced. * The experimental evaluation is a bit terse; I would like to see the performance of this learning rule on a broader range of tasks (especially a non-toy problem with some temporal structure).

Questions

* Line 150: learning is duplicated * Figure 2(a): The inset mentioned in the caption seems to be missing

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

All of the concerns I had were addressed thoroughly in the discussion section.

Reviewer z6yQ5/10 · confidence 4/52023-07-06

Summary

This paper adds recurrent inhibition to each layer of a DNN architecture in order to facilitate a more biologically-plausible form of credit assignment. It shows how this circuit can explain some of the features of plasticity found in vitro and that DNNs with this circuit can learn to perform simple visual tasks.

Strengths

The microcircuit is biologically motivated The insights provided into how the artificial experimental conditions in plasticity studies lead to specific results is helpful

Weaknesses

The motivation and innovation was not entirely clear to me. The background is focused on how credit assignment requires distant neurons to interact, yet the problem of how information reaches each layer in the network is not what is tackled here. There is also discussion of the microcircuit "decoding" the credit signal, but it seems like the credit signal is fairly directly given to the layer, and so the microcircuit is more of a relay than a decoder. One of the main results seems to be that a linear approximation to the inverse activation function can make the weight update rule slightly more biologically realistic and still works fairly well. This is fine, though not a very impactful result. I found some of the results descriptions confusing (see below)

Questions

Why is there no weight matrix for the recurrent connections? Can the authors elaborate on how their formulation supports stability? For example, I didn't fully understand this claim "On the other hand, our model suggests that a rapid compensatory mechanism could be implemented as a combination of recurrent inhibitory microcircuits and a linear inhibitory threshold in the synaptic plasticity rule" "Experiments on excitatory plasticity in vitro are commonly performed in the presence of Tetrodotoxin, a sodium channel blocker, to minimize interference of inhibitory activity" Doesn't TTX block all activity, not just inhibitory? "Specifically, we block recurrent inhibition " From the diagram it looks like you block recurrent excitation, not inhibition (There would be no point in varying the activity of the inhibitory neuron if its outward connections were blocked). Is "in the absence of any inhibitory input" supposed to say "absence of input to the inhibitory cells" ? It's not clear to me why the "postsynaptic" term in equation 9 drops out when the E->I connection is dropped "we manually set the feedback weights to the transposed network Jacobian at each time step" does this mean that the feedback is different at each timestep? That is not very biologically-plausible.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

2 fair

Presentation

3 good

Contribution

2 fair

Limitations

Already mentioned above

Reviewer TQSB2023-08-10

Response to rebuttal

Thank you for your very detailed feedback. Though I still believe the results in this paper are incremental relative to previous work, and that this is the principal weakness of the paper, you have done a lot to convince me of the rigor of your analyses. Adding these points (especially your additional figures) will certainly increase the quality of the paper. I will maintain my score, but will increase my confidence (score: 6; conf. 5).

Reviewer A5Rt2023-08-12

Thank you to the authors for the detailed and informative response. I'm impressed with the work the authors have done in the rebuttal and my main concerns have been addressed. Overall, I would not call this model a massive leap forward from the Meulemans et al. works, but I think the authors do make sufficiently novel and interesting predictions for neuroscience, a field in which gains are typically incremental. I will upgrade my score to 7.

Reviewer EPnb2023-08-13

Thank you for the detailed response, and especially the additional simulations! Although I do maintain that one-to-one connectivity prevents these results from being a true breakthrough in biologically-plausible backprop, I believe this is a good paper that makes tangible progress and stand by my original score.

Reviewer z6yQ2023-08-14

I appreciate the clarifications and corrections from the authors. The intended impact of the work is now clearer. I will increase my score by 1.

Authorsrebuttal2023-08-21

We would like to thank the reviewers for their thorough review and constructive feedback. We appreciate the time dedicated to our paper and look forward to improving it based on the provided suggestions.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC