On the Limitations and Capabilities of Position Embeddings for Length Generalization

In Transformers, Position Embeddings (PEs) significantly influence Length Generalization (LG) performance, yet their fundamental role remains unclear. In this work, we identify two fundamental limitations of PEs in achieving LG: the inability to acquire new operators beyond training data, and the inability to deal with inconsistent computational roles of the same positions across scales. To formalize these limitations, we introduce computational representation complexity which measures the number of distinct unit operators required for a task, and the notion of canonicality, which characterizes operator-position consistency as the scale varies. We prove that PEs cannot achieve LG for mappings whose computational representation complexity increases with scale, nor for mappings that are non-canonical. We further show that PEs enable LG when these two failure modes are absent. In this case, LG can be achieved with a PE whose positional relation function characterizes the computational representation. To enhance LG for non-canonical mappings, we introduce scale hint, allowing flexible instance scaling. We also propose a learning-based position embedding framework that automatically learns positional relations. Our work provides theoretical insights and practical strategies for improving LG in Transformers.

Paper

Similar papers

© 2026 NYSGPT2525 LLC