Understanding the Mechanics of SPIGOT: Surrogate Gradients for Latent Structure Learning

Latent structure models are a powerful tool for modeling language data: they\ncan mitigate the error propagation and annotation bottleneck in pipeline\nsystems, while simultaneously uncovering linguistic insights about the data.\nOne challenge with end-to-end training of these models is the argmax operation,\nwhich has null gradient. In this paper, we focus on surrogate gradients, a\npopular strategy to deal with this problem. We explore latent structure\nlearning through the angle of pulling back the downstream learning objective.\nIn this paradigm, we discover a principled motivation for both the\nstraight-through estimator (STE) as well as the recently-proposed SPIGOT - a\nvariant of STE for structured models. Our perspective leads to new algorithms\nin the same family. We empirically compare the known and the novel pulled-back\nestimators against the popular alternatives, yielding new insight for\npractitioners and revealing intriguing failure cases.\n

Paper

References (61)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC