**Sequence Lengths**
In LD4LG, we perform diffusion in a compressed latent space (e.g. dimension 32 x 64). We can compress texts of different lengths into this latent space and then rely on the decoder’s ability to generate text sequences of different sizes when conditioned on the diffused latent code. We present some statistics for the lengths of 1000 random validation samples across various datasets we examined. We use simple whitespace tokenization to compute these statistics to be tokenizer-agnostic. The numbers are therefore low compared to the number of BPE tokens for the various models. We also compute the Pearson’s r Correlation between the input length and the length of the autoencoder reconstruction.
| Dataset | Mean Length | StDev Length | Min Length | Max Length | Pearson's r Correlation (Input Length vs. Reconstruction Length) |
|-|-|-|-|-|-|
| ROCStories | 43.5 | 7.7 | 20 | 58 | > 0.99 |
| AG News | 30.0 | 8.3 | 8 | 56 | > 0.99 |
| WMT14-En | 18.8 | 10.9 | 1 | 77 | > 0.99 |
We observe that all datasets have a variety of lengths and that the input and output lengths have near-perfect correlation (> 0.99). The autoencoder therefore effectively reconstructs input language of a variety of lengths from the fixed-size latent space. The near-perfect reconstruction performance reported in our submission (Table 1) also demonstrate this.
We also present some examples of our autoencoder’s reconstructions for validation inputs of varying lengths from the WMT-14 English dataset.
| Input | Reconstruction |
|-|-|
| That pleases me. | That pleases me. |
| For almost ten years the choir has been practising songs in this foreign,'soft' language, and now and then they bring them back to where they originally came from: the South of Africa. | For almost ten years the choir has been practising songs in this foreign,'soft' language, and now and then they bring them back to where they originally came from: the South of Africa. |
| Then she tells of her husband who was in the army. | Then she tells of her husband who was in the army. |
| Amid the usual futile arguments over who started it, scores of buildings have been reduced to rubble; more than 140 Palestinians, most of them civilians, and six Israelis have been killed; and, for the first time, missiles from Gaza have landed near Tel Aviv, Israel's metropolis, and the holy city of Jerusalem. | Amid the usual futile arguments over who started it, scores of buildings have been reduced to rubble; more than 140 Palestinians, most of them civilians, and six Israelis have been killed; and, for the first time, missiles from Gaza have landed near Tel Aviv, Israel's metropolis, and the holy city of Jerusalem. |
| Whether at a physical comfort, emotional or spiritual level. | Whether at a physical comfort, emotional or spiritual level. |
| It starts by the Antonia Fortress - Praetorium - where the judgement took place, and brings us along the streets of the Old Town to the Church of the Holy Sepulchre on Golgotha - the place of the crucifixion, Stone of Unction and the place of Jesus' burial. | It starts by the Antonia Fortress - Praetorium - where the judgement took place, and brings us along the streets of the Old Town to the Church of the Holy Sepulchre on Golgotha - the place of the crucifixion, Stone of Unction and the place of Jesus' burial. |
We would like to emphasize that any limitations that our model has in terms of modeling datasets of different lengths will be shared by other diffusion and autoregressive models. Other language diffusion models, such as Diffusion-LM, generate instances of varying length by always generating a max-length sequence with a variable number of [PAD] tokens. To generate “It pleases me.”, for instance, Diffusion-LM would generate:
“[BOS] It pleases me. [EOS] [PAD] [PAD] [PAD] [PAD] [PAD]... [PAD]”.
Our hybrid approach represents an improvement over such methods.
Autoregressive models often fail to extrapolate to longer sequence lengths than those observed during training [1]. This limitation would be shared by our framework, but it is not unique to our approach. While overcoming this limitation in autoregressive models is an active area of research [2], such advances could potentially be incorporated into our hybrid framework.
[1] Kazemnejad, Amirhossein, et al. "The Impact of Positional Encoding on Length Generalization in Transformers." (2023).
[2] Press, Ofir, Noah A. Smith, and Mike Lewis. "Train short, test long: Attention with linear biases enables input length extrapolation." (2021).