The Corpus is the Moat: Retrieval-Augmented Generation and the Strategic Repositioning of Libraries, Archives, Journals and Book Publishers

Purpose — This paper argues that retrieval-augmented generation (RAG), the AI architecture that grounds language-model output in an external corpus and returns citations, has an unusually favourable alignment with the historical commitments and professional assets of memory institutions. It sets out the case that libraries, archives, scholarly journals and book publishers are structurally positioned not as victims of generative AI, but as its supply side. Design/methodology/approach — The paper is a conceptual and integrative review. It synthesises the technical RAG literature, empirical evaluations of RAG-based scholarly search tools, and current industry evidence on AI-rights licensing, integrated with a structured reading of Global South infrastructure and language-resource studies. The argument is grounded in Lewis et al.'s (2020) original formulation of RAG, extended through Bevara et al.'s (2025) academic-library-specific treatment, and stress-tested against a single-institution case in a Nigerian central bank library. Findings — Three convergent findings emerge. First, the two open problems that RAG was designed to address — provenance and knowledge currency — are the two founding commitments of the information professions, so the architecture and the profession share a common object. Second, model capability is commoditising while the scarce input has shifted to the corpus: curated, deduplicated, described and legally usable collections have become priced assets, with major publishers already recording substantial AI-licensing revenue. Third, the constraints usually treated as disqualifying for AI adoption in the Global South — intermittent power, expensive bandwidth, foreign-currency pricing, and language exclusion — are in fact more favourable to RAG than to model-training approaches, provided the design privileges retrieval, small local models, and multilingual corpus-building. Research limitations/implications — The paper is conceptual and integrative rather than empirical. The single-institution case is illustrative, not generalisable, and is anonymised at the institution's request. The rapidly shifting technical landscape means that specific vendor claims cited here may age within twelve months. Practical implications — A staged, five-phase adoption roadmap is proposed, together with a set of preconditions that determine whether a RAG deployment will succeed: retrieval quality as the ceiling on system quality; the necessity of measured faithfulness; a designed abstention policy; rights hygiene; and provenance as a non-negotiable output. Design principles for constrained environments — offline-first architectures, retrieval-heavy investment, small local models, and consortial procurement — are articulated as an antidote to the implicit assumption of reliable infrastructure that pervades the current literature. Social implications — If the Global South consumes RAG systems built exclusively on Northern corpora, it will import both the knowledge and the silences of those corpora. If it builds on its own collections — including in Hausa, Yoruba, Igbo and other under-represented African languages — it places its scholarship into the global retrieval layer for the first time. Originality/value — The paper offers, to the authors' knowledge, the first cross-sector treatment of RAG as a single architectural problem confronting libraries, archives, journals and book publishers together; the first analytical connection between the RAG rights market and the classical rights-and-licensing infrastructure of publishing; and the first substantial Global South analysis of the architecture from inside its infrastructural constraints.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC