Towards training digitally-tied analog blocks via hybrid gradient computation

Power efficiency is plateauing in the standard digital electronics realm such that novel hardware, models, and algorithms are needed to reduce the costs of AI training. The combination of energy-based analog circuits and the Equilibrium Propagation (EP) algorithm constitutes one compelling alternative compute paradigm for gradient-based optimization of neural nets. Existing analog hardware accelerators, however, typically incorporate digital circuitry to sustain auxiliary non-weight-stationary operations, mitigate analog device imperfections, and leverage existing digital accelerators.This heterogeneous hardware approach calls for a new theoretical model building block. In this work, we introduce Feedforward-tied Energy-based Models (ff-EBMs), a hybrid model comprising feedforward and energy-based blocks accounting for digital and analog circuits. We derive a novel algorithm to compute gradients end-to-end in ff-EBMs by backpropagating and"eq-propagating"through feedforward and energy-based parts respectively, enabling EP to be applied to much more flexible and realistic architectures. We experimentally demonstrate the effectiveness of the proposed approach on ff-EBMs where Deep Hopfield Networks (DHNs) are used as energy-based blocks. We first show that a standard DHN can be arbitrarily split into any uniform size while maintaining performance. We then train ff-EBMs on ImageNet32 where we establish new SOTA performance in the EP literature (46 top-1 %). Our approach offers a principled, scalable, and incremental roadmap to gradually integrate self-trainable analog computational primitives into existing digital accelerators.

Paper

Similar papers

Peer review

Reviewer pevw6/10 · confidence 3/52024-06-21

Summary

State-of-the-art (SOTA) analog hardware accelerators consist of both analog and digital components supporting major and auxiliary operations. Moreover, they typically suffer from device imperfections on their analog parts. In this paper, the authors propose feedforward-tied energy-based models (ff-EBMs) for digital and analog circuits. ff-EBMs compute gradients by both backpropagating on feedforward parts and eq-propagating on energy-based parts. In this paper, ff-EBMs use a Deep Hopfield Network (DHN) as one example, which can be arbitrarily partitioned into uniform size. ff-EBMs achieves a SOTA performance on ImageNet32.

Strengths

1. The paper works on an important and interesting problem. 2. The paper flows well.

Weaknesses

1. The paper does NOT consider a detailed analog computing device model. It will be interesting to see how significant variations on the analog devices the ff-EBMs can tolerate. And what is the relationship between the convergence speed and the device variations? 2. The paper does not include energy-related data. What is the energy saving for the energy-based models? What is the energy saving of skipping the accurate gradient computations?

Questions

Please comment on the points in the weakness section.

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

no limitations.

Reviewer kDhg7/10 · confidence 4/52024-07-04

Summary

This paper discusses a novel approach to improving the power efficiency of AI training by integrating analog and digital hardware. The authors introduce a hybrid model called Feedforward-tied Energy-based Models (ff-EBMs) that combines feedforward neural networks with energy-based models (EBMs). This model aims to leverage the benefits of both analog and digital systems to create a more efficient computational framework for neural network training. The paper introduces a new model, ff-EBMs, which integrates feedforward and EB modules. These models are designed to perform end-to-end gradient computation, combining backpropagation through feedforward blocks and "eq-propagation" through EB blocks. And bilevel and multi-level optimization are concepts used to frame the learning process of the Feedforward-tied Energy-based Models (ff-EBMs). The paper promises to demonstrate the effectiveness of the proposed approach experimentally, using DHNs as EB blocks and achieving new state-of-the-art performance on the ImageNet32 dataset.

Strengths

- The paper presents a robust framework that leverages the strengths of bilevel and multi-level optimization to address the challenges of training neural networks in a hybrid digital-analog setting. One of the significant strengths lies in the rigorous mathematical proofs provided for the bilevel optimization inherent to energy-based models (EBMs). The authors extend this foundation to a multi-level optimization approach, which is not only theoretically sound but also practically applicable to the training of Feedforward-tied Energy-based Models (ff-EBMs). - The bilevel optimization is solidly established with the inner problem accurately reflecting the equilibrium state search, crucial for EBMs, while the outer problem adeptly encapsulates the parameter adjustment to minimize the cost function. This nested structure is then expanded into a multi-level optimization problem that inherently captures the complexity of training ff-EBMs, where each level represents an optimization challenge within the model's architecture. - The theoretical robustness is complemented by a well-thought-out algorithm that intertwines backpropagation and equilibrium propagation, offering a practical solution to the end-to-end training of these hybrid models. This algorithm is not just a theoretical construct; it is underpinned by a series of proofs that validate its effectiveness. The experimental results are particularly compelling, as they demonstrate the effectiveness of the proposed approach on ff-EBMs, achieving state-of-the-art performance on the ImageNet32 dataset.

Weaknesses

While the paper introduces an innovative approach to hybrid digital-analog neural network training, there are a few areas where it shows some limitations. - The paper could benefit from clearer articulation regarding the energy efficiency claims, providing more detailed comparisons with existing systems to substantiate these claims. - Additionally, the reliance on simple datasets, ImageNet32, for experimental validation, while common, might limit the generalizability of the findings. A broader range of datasets with varying characteristics could strengthen the paper's conclusions. The choice of ImageNet32 should be justified with respect to its relevance to the research goals and its limitations for testing the model's robustness. - Lastly, while the authors acknowledge the need for further research, a more explicit discussion on the current limitations and a roadmap for future work would provide a more comprehensive view of the research's trajectory.

Questions

In the proposed implicit BP-EP chaining algorithm, is it necessary to ensure the satisfaction of the first-order optimality conditions derived from the implicit theory, or can the algorithm be robust and effective without strictly guaranteeing these conditions?

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

A notable limitation of our study is the constrained scope of network architectures and datasets utilized for validation. While the implicit BP-EP chaining algorithm has shown success with Deep Hopfield Networks on ImageNet32, its performance on the standard ImageNet and transformer-based models, which are more complex and widely used, is yet to be established.

Reviewer 9JF87/10 · confidence 4/52024-07-12

Summary

This paper proposes a new building model block with analog forward circuits and energy-based blocks built on a digital-analog hybrid setup. A novel algorithm is further proposed to train the new block model. Experiments show the SOTA accuracy in the EP literature.

Strengths

* This paper achieves SOTA accuracy compared to recent literature. * A solid optimization method is proposed for the hybrid block.

Weaknesses

* The paper's motivation is vague to me, as I am new to this topic. Why must we combine an analog forward circuit with a digital one for EP-based training? For the analog accelerator, we can use several other methods to do training, e.g., forward-forward only and zeroth-order optimization. What are the benefits of incorporating an energy-based model?

Questions

N/A

Rating

7

Confidence

4

Soundness

3

Presentation

2

Contribution

3

Limitations

N/A

Reviewer k7Rs5/10 · confidence 3/52024-07-15

Summary

This paper presents Feedforward-tied Energy-based Models (ff-EBMs), a hybrid model that integrates feedforward and energy-based components, accounting for both digital and analog circuits. A novel algorithm is proposed to compute gradients end-to-end in ff-EBMs by backpropagating and "eq-propagating" through feedforward and energy-based sections, respectively, allowing EP to be applied to more flexible and realistic architectures. It has been shown that ff-EBMs can be trained on ImageNet32, achieving new state-of-the-art performance in the EP literature with a top-1 accuracy of 46%.

Strengths

-- The proposed ff-EBMs as high-level models of mixed precision systems, where the inference pathway is composed of feedforward and EB modules, is interesting and novel. Specially, gradients computations as an end-to-end backpropagation through feedforward blocks and “eq-propagating” through EB blocks. -- The results are also encouraging specially on CIFAR datasets. -- The paper is easy to read and understand.

Weaknesses

-- The primary limitation of this work is its accuracy performance on the ImageNet dataset. Although the paper has set a new state-of-the-art accuracy, it still lags behind the results achieved by Transformers and CNNs. There is a lack of compelling reasons to use this method given its comparatively lower performance. -- Additionally, it is crucial to measure the energy consumption of this training method and compare it with traditional methods. How much energy savings does your method offer?

Questions

See weaknesses.

Rating

5

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

NA

Reviewer Jt5z5/10 · confidence 3/52024-07-31

Summary

Analog in-memory computing is gaining traction as an energy-efficient platform for deep learning. However, fully analog-based accelerators are challenging to construct, necessitating a training solution for digital-analog hybrid accelerators. This paper introduces Feedforward-tied Energy-based Models (ff-EBMs), a hybrid model that integrates feedforward components, typical in digital systems, with energy-based blocks suitable for analog circuits. An algorithm is derived to compute gradients end-to-end for training ff-EBMs. Experimental results show that ff-EBMs achieve superior accuracy on ImageNet32 classification compared to conventional equilibrium propagation literature.

Strengths

1. This paper developed training algorithms specifically for digital-analog hybrid computing platforms, addressing a realistic computing scenario that leverages analog in-memory computing.

Weaknesses

1. The paper lacks examples of situations where ff-EBM is needed. It does not adequately explain which parts of the overall system are digital and which are analog when using ff-EBM. Additionally, it does not clarify whether the existing training techniques can be used in these scenarios or what advantages ff-EBM offers compared to traditional methods.

Questions

As I understand it, analog in-memory computing naturally finds the node voltage that minimizes the energy function through Kirchhoff's laws, classifying it as an EP problem. Therefore, if we disregard the natural minimization of the energy function, the training process with analog in-memory computing appears very similar to the conventional backpropagation-based training procedure. Hence, the derived algorithm for calculating end-to-end gradients of the ff-EBM seems trivial. In this context, I don't see the difference between this research and previous studies that have conducted training on analog in-memory systems using EP or backpropagation. Could you explain the differences between the previous approaches and the proposed approach in more detail? Providing examples of situations where ff-EBM is needed, as discussed in the Weakness section of this review, would be helpful.

Rating

5

Confidence

3

Soundness

3

Presentation

1

Contribution

3

Limitations

Please check the weakness.

Reviewer Jt5z2024-08-08

Additional questions on the novelty of this paper.

Thank you very much for your careful and detailed response. Figure 1 and Figure 2 in the attached PDF file clearly illustrate the hardware set-up and the proposed layer design. I acknowledge that this paper is pioneering in addressing the training algorithm for hardware composed of multiple analog chips, and I agree that this is a realistic configuration. However, I still have a few questions regarding the novelty of this paper. The ff-EBM model architecture comprises a chain of the FF module and the EB module. Given that the training algorithms for the FF module and EB module are established conventions (backpropagation and equilibrium propagation, respectively), I find it challenging to identify the difficulty in designing a training algorithm for ff-EBM. Since the training algorithms for the FF and EB modules are predetermined, the primary task for ff-EBM is to ensure the proper passing of gradients between the FF and EB modules. Nonetheless, passing gradients between these modules appears straightforward by following the chain rule. Therefore, I would like the authors to highlight any algorithmic challenges related to configuring the FF and EB modules within a single network architecture. More specifically, I would appreciate it if the authors could provide more detailed explanations on the challenges involved in passing gradients between the FF and EB modules and how these challenges are addressed.

Authorsrebuttal2024-08-08

Answering additional questions on the novelty of the paper

Dear Reviewer Jt5z, We thank you very much for engaging promptly in this discussion, and are happy indeed to read that the PDF file brought some satisfying clarifications about the relevance of the problem tackled. If we understand you correctly, since BP and EP are well established algorithms in their own right, chaining them inside a given architecture appears straightforward “by following the chain rule”. In fact, we are very pleased that you have drawn attention to this issue, as we are increasingly recognizing that some of the terminology we use could prove misleading, and to some extent confusing for readers. Therefore, as requested, we highlight below the specific “algorithmic challenges” pertaining to this chaining, “how [they] are addressed” and more broadly the “novelty of this paper”. - **EP-BP chaining is intuitive, but not trivial to derive**. You write that chaining EP and BP inside a given architecture “appears straightforward by following the chain rule”. We would like to focus attention on this question: what is meant by “following the chain-rule” in the context of ff-EBMs? + Traditionally, “chain-rule” refers to gradient chaining inside *feedforward models*. Namely, let us assume a computational graph of the form $s^1 = F(x) \to s^2 = G(s^1)$. Assuming an error signal $\delta^2 = \partial_{s^2} L$, then the “chain rule” prescribes that the error signal at $x$ reads $\partial_{x}L = \partial_x F(x)^\top \cdot \partial_{s^1}G(s^1)^\top \cdot \delta^2$. + Now let us assume instead the following “hybrid” computational path: $s^1 = F(x) \to s^2: \nabla_{s^2}E(s^2, s^1) = 0$. Given that $s^2$ in this case is an *implicit* function of $s^1$, i.e. there is no explicit G mapping as before between $s^1$ and $s^2$, how would one directly apply the above “chain-rule” here? To put it differently, how do we rigorously route error signals backward through this computational graph? This is the first “challenge” we addressed. + Hence the need to derive gradient chaining inside ff-EBMs *rigorously, from first principles* by: 1) stating the learning problem as a multilevel constrained optimization problem (Eq. 8), 2) writing the associated Lagrangian (Eq. 21), 3) solving for the associated KKT conditions (Eqs. 22-34) for the primal variables (i.e. the steady states of the blocks) and associated Lagrangian multipliers (i.e. the error signals inside blocks). Therefore, this is how “we addressed” the above challenge. We emphasize that **this derivation** (namely Theorem 3.1 about the rigorous chaining of EP and BP gradients) and the **resulting explicit and implicit chaining algorithms** (Alg. 4 inside Appendix A.3, Alg. 2 in the main) **are novel**. Also, note from the above that our algorithm comes into an *explicit* and *implicit* variant (see paragraph “Proposed algorithm: implicit BP-EP chaining” L.187 and Lemma A.4), the latter appearing as a “pure” EP implementation. As such, the fact that EP-BP chaining can be cast into a *non-trivial generalization of EP*, as appearing in Alg. 2 and Lemma A.4, is not self-evident neither. - **The experimental demonstration of this algorithm is also novel**. Finally, to further highlight the “novelty of this paper”, **our algorithm was never tested in practice before this work** and having it succeed on ImageNet32 came with lots of “challenges” as well, the most important one being the simulation time. We addressed this problem in two ways: + As explained in the global rebuttal and in greater details inside Reviewer k7Rs’s rebuttal, splitting an EBM into several EBM blocks tied by feedforward modules results in an architecture that is not only more hardware realistic, but also **easier to simulate**. Indeed, instead of the superlinear scaling of the simulation time with respect to the number of layers observed in past EP works, our approach guarantees, by construction, that this **simulation time scales linearly with the number of blocks**, each of these blocks converging much faster than the full EBM counterpart. See our new table of results inside our PDF attached to this rebuttal. While the ff-EBMs trained are still relatively shallow, **they are twice as deeper as the deepest EBM trained by EP** in most recent related works. + Finally, using *Gaussian Orthogonal Ensembles* (GOE) to initialize weights inside EB blocks was instrumental in having ff-EBM training experiments work. Finally, **we would like to propose the addition of a few sentences to our introduction highlighting this novelty**, and particularly the degree to which our algorithm is derived by exploiting an intimate theoretical connection between energy based learning and **implicit differentiation**, rather than literal "backprop" as applied in standard feedforward nets where the standard "chain rule" applies. Hence our claim that this contribution belongs to the realm of EP, and *implicit learning* more broadly. Again, we thank you for drawing our attention to this.

Reviewer Jt5z2024-08-11

Thank you for your detailed explanation. Your response has helped me better understand the proposed work, and I have made every effort to assess its value. Firstly, I am increasing my score to 5, as the novelty of your work is now clear to me. From my understanding, EP is designed to be applicable to any network architecture, including those with feedforward blocks, as the minimization of energy can be aligned with the minimization of the objective function (as shown in Figure 1 and Chapter 3 of the BP paper [1]). Therefore, I still view Eq. (8) as a straightforward integration of FF and EBM. However, I acknowledge that the gradient calculation starting from Eq. (8) is not trivial, and the detailed derivation of these gradients is a key novelty of your work. While I believe this work is highly significant in the field of analog-based AI accelerator systems, I am hesitant to raise my score further because the presentation of your work could be improved to better highlight its true value. In my opinion, the design of networks for analog-based computing is heavily constrained by hardware considerations, as there are significant challenges in scaling fully analog systems. Analog computing is known for its energy efficiency compared to digital computing, while digital computing offers greater scalability. To build an efficient yet scalable AI acceleration system, a hybrid design is essential. In this context, I believe the true value of your work lies in extending the scalability of EP-based models for analog computing by integrating FF modules. For this reason, I think that when adopting ff-EBM, the key concern should be scalability. The scalability of ff-EBM, as compared to using EBM alone, is a critical point that should be thoroughly explored. For example, with EBM alone, only a single analog-based unit in Figure 1 of attached PDF could be used for a single network. Obviously, ff-EBM should be much better than EBM, but I believe this paper needs detailed discussion on the accuracy of this EBM-based model and how it compares to the accuracy of a larger model designed with ff-EBM that fully utilizes the entire system depicted in Figure 1 of attached PDF, as this is an important aspect of this work. [1] B. Scellier and Y. Bengio. Equilibrium propagation: Bridging the gap between energy-based models and backpropagation. Frontiers in computational neuroscience, 11:24, 2017.

Authorsrebuttal2024-08-12

Clarifying Scellier-Bengio's EP paper & scalability of ff-EBMs

We are very grateful to Reviewer Jt5z for spending time to understand the value of our paper, subsequently increasing our score and engaging in this discussion which will tremendously benefit the presentation of our work. Thank you so much! In the light of Reviewer’s Jt5z last answer, we would like to clarify some essential points they raised: - *"EP is designed to be applicable to any network architecture, including those with feedforward blocks"*. As mentioned in Section 2.3 of our paper, **EP only applies to energy-based models**. Fig. 1 of the seminal EP paper [1] is indeed misleading: when writing “Equilibrium Propagation applies to any architecture”, one should understand “any architecture **topology** so long as it derives from an energy function”. Indeed, as indicated by the title of the section 3 of this paper, EP really is a “Machine Learning Framework **for Energy-Based models**”: their Figure 1 is only meant to emphasize that EP applies to *any* energy-based models, not necessarily *layered* energy-based models. In this context: **layered does not meant feedforward**. As Scellier & Bengio write themselves: “*In particular, the [EP learning rule] holds for any architecture and **not just a layered architecture** (Figure 1) like the one considered by Bengio and Fischer (2015)* [which is also an EB model]”. - *"Therefore, I still view Eq. (8) as a straightforward integration of FF and EBM"*. Given the clarification above, this conclusion may no longer hold, especially when noticing that Eqs.17-18 of the seminal EP paper [1] (in the section 3 mentioned by Reviewer Jt5z), i.e. the bilevel program the EP algorithm solves, **is an explicit particular case of the Eq. 8 of our paper**, i.e. the multilevel program that our algorithm solves. - *"I think that when adopting ff-EBM, the key concern should be scalability"*. We totally agree! Investigating the scalability of ff-EBM training by our algorithm on deeper ff-EBMs, more complex tasks and exploring new datasets and architectures is part of our research roadmap. See our detailed answer to Reviewer kDhg and associated Fig. 3 inside the PDF. - *"The scalability of ff-EBM, as compared to using EBM alone, is a critical point that should be thoroughly explored"*. We are happy Reviewer Jt5z mentions the importance of this comparison since our “splitting experiment” (Section 4.3) goes exactly in this direction. This experiment reveals that a single EBM block performs comparably to an ff-EBM **of equal depth** with various block sizes. [1] Scellier, B., & Bengio, Y. (2017). Equilibrium propagation: Bridging the gap between energy-based models and backpropagation. Frontiers in computational neuroscience, 11, 24.

Reviewer 9JF82024-08-12

Thank you for your rebuttal

Thank you for your rebuttal. First, thank you for your new Figure 2, which makes your problem statement much clearer to me, while your previous introduction is too complex and hard to get your point. So, if I understand correctly, **you propose to make the analog part to an energy model and set the digital part to the normal feedforward layer. Then, you derive the automatic differentiation method to train this hybrid setup**. So, I have several key follow-up questions to help me decide whether to raise/lower my score. * Which is the main source of the accuracy improvement? The hybrid setup is where the digital feedforward improves accuracy, or the beyond zeroth-order backpropagation algorithm is used. * The reason I ask ZO/FF is not due to the accuracy; it is because ZO/FF is hardware-friendly for analog circuits where you cannot use backpropagation to do on-chip training. Therefore, ZO/FF is practical for analog hardware requiring forward passes. Is your proposed method also friendly? How can you obtain the right gradient through the analog circuit? Please discuss the real hardware implementation for your method, especially how to obtain the gradient. thank you so much.

Authorsrebuttal2024-08-12

Source of improvement & hardware implementation of our algorithm

We are glad that Reviewer 9JF8 appreciated our rebuttal, and are indeed happy to clarify their questions. - *you propose to make the analog part to an energy model and set the digital part to the normal feedforward layer*. Yes: analog and digital parts account for energy-based and feedforward models respectively. - *you derive the automatic differentiation method to train this hybrid setup*. To be certain to be on the same page as Reviewer 9JF8, we slightly reformulate this sentence: + we apply the *Lagrangian method* to derive optimality conditions, i.e. “KKT” conditions, to “train this hybrid setup” (Appendix A.2). + the application of this method yields an algorithm which is itself a **hybrid** differentiation method which uses standard backprop (i.e. “automatic differentiation”) inside feedforward / digital parts and **equilibrium propagation inside energy-based / analog parts**. Therefore, our algorithm isn’t “pure” automatic differentiation end-to-end, but *only within feedforward parts of the model*. - *Which is the main source of the accuracy improvement?* The main source of accuracy improvement, *with respect to the past EP works*, is simply the *depth* of the model trained: the ff-EBMs trained here are twice as deep as the largest EBMs trained in previous EP works. Reviewer 9JF8’s question may then translate to: why can you train deeper networks? The most important reason, which is highlighted in Table 1 of our PDF and in our rebuttal to k7Rs, is that **ff-EBMs are faster to simulate than their fully EBM counterparts** (up to 4 times faster per Table 1). Instead of the superlinear scaling of the convergence time of DHNs (with respect to the number of layers) empirically observed in past EP works, the convergence time ff-EBMs made up of DHNs are EB blocks are guaranteed, by construction, to scale linearly with the number of blocks, the convergence time of a single block decreasing with its size. On the algorithmic side: since i) our algorithm performs on par with end-to-end automatic differentiation on the same ff-EBMs on the one hand, and that ii) ZO techniques are expected to perform less well than automatic differentiation on equivalent architectures (as highlighted in our rebuttal to Reviewer 9JF8), we can conjecture that, on *equal ff-EBMs*, our algorithm may likely perform better than ZO *when applied end-to-end*. - *“Is your proposed method hardware-friendly? Please discuss the hardware implementation of your method”*. First, as highlighted in our rebuttal, our method extends EP-based training to scaled hardware systems which: i) may not fit a single analog core, ii) may still require operations which cannot be supported on analog hardware and may instead require digital, high precision hardware. Therefore, Fig. 1 of the PDF attached to our rebuttal is a plausible depiction of the hardware implementation of our method and taking these constraints into account, our algorithm is **hardware plausible**. Second, we insist that EP-based training inside each of the analog / energy-based block inherits the “hardware-friendly” / FF-like features of EP training: in our algorithm, **gradients inside analog / energy-based blocks are computed using only forward passes / relaxations to equilibrium**. A plausible hardware implementation of this analog blocks are *deep resistive networks* [1, 2]. Lastly, feedforward blocks, which are maintained in digital, could also directly leverage quantization algorithms [3] and even **ZO algorithms** (as mentioned in L.290-291 of our paper) to facilitate their implementation on memory-constrained, low-power hardware. To summarize: our algorithm makes the best of analog and digital worlds, by: i) being hardware plausible at the system level (Fig .1 of PDF), ii) preserving the “hardware-friendliness” of FF-like learning inside EB/analog blocks, iii) possessing the ability to leverage any quantization algorithm inside feedforward blocks and iv) possessing the ability **to apply ZO algorithms** instead of backprop (as currently done with our algorithm) **inside feedforward blocks**. Finally, we would like to thank reviewer 9JF8 for drawing our attention to some critical points where our paper could better communicate some of the high-level motivations and details of our algorithm. In addition to what was proposed in our original rebuttal, we would like to propose adding a pseudo algorithm in the appendix showing how ZO could be applied within feedforward blocks (instead of backprop currently) such that gradients would be computed **everywhere with forward passes only**, i.e. both inside analog blocks (by EP) and feedforward blocks (by ZO) (as first suggested in L. 290-291 of our paper). [1] Kendall et al (2020). Training end-to-end analog neural networks with equilibrium propagation. [2] Scellier, B. (2024). A Fast Algorithm to Simulate Nonlinear Resistive Networks. ICML 2024. [3] Lin et al (2022). On-device training under 256kb memory. NeurIPS 2022

Authorsrebuttal2024-08-13

Dear Reviewer 9JF8, thank you so much for taking the time to engage in a discussion with us and for raising our score! Your remarks will greatly help improve our manuscript for readers interested in ZO optimization. Thank you again!

Reviewer kDhg2024-08-13

Respond to rebuttal.

Thank you for your reply. The reviewer has no additional questions.

Program Chairsdecision2024-09-25

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC