Existing research on lexical bundles (e.g. Biber 2009) suggests that 'literate' registers such as academic prose contain shorter bundles than 'oral' registers like conversation. This phenomenon is connected to formulaicity: in more formulaic and redundant registers, such as conversation, it is more economical for speakers to retain and reuse longer sequences. This principle, however, is not economical in registers that contain a lot of variation, such as very formal academic articles. We propose that this connection affects the formation of idiolectal lexical bundles, those n-grams that identify a single individual. That is, idiolectical lexical bundles may be longer in more formulaic registers compared to less formulaic ones. We analysed six corpora spanning registers from academic prose to online chats, measuring the degree of 'literacy' of a register using the Formality score (Heylighen and Dewaele 2002). Formulaicity was measured by analysing document samples into n-grams of varying lengths and computing the Jaccard coefficient. We then ran a set of authorship verification experiments using N-gram Tracing (Grieve et al., 2018) to identify the optimal n-gram length for each register. Results reveal a systematic negative relationship between formality and formulaicity: more formal corpora exhibited lower levels of n-gram repetition than more informal registers. Furthermore, shorter n-grams (n=1 to n=2) achieved significantly higher accuracy for the most formal corpora, while longer n-grams (n=3 to n=6) performed comparably to shorter ones in less formal corpora.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex