Large Language Models (LLMs) represent a transformative advancement in natural language processing (NLP), building upon foundational Language Models (LMs) to achieve human-like language understanding and generation through massive scale and sophisticated architectures. This paper provides a comprehensive overview from a computer science lens, defining LMs and LLMs, dissecting the Transformer-based architecture central to LLMs, exploring their functionalities, and contrasting them with traditional LMs. Key components like self-attention and positional encodings are detailed with mathematical formulations, while a glossary and references ensure accessibility. By highlighting scaling laws and emergent abilities, we underscore LLMs' role in enabling zero-shot learning and multimodal applications, alongside challenges like computational efficiency and ethical considerations. This analysis serves as a primer for researchers and practitioners who are looking to navigate the evolution of AI-driven language technologies while offering a systematic framework to compare LLM architectures and emerging behaviors.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex