Exploring Open-Weight Foundation Models: A Systematic Review from Transformers to Multimodal AI
Foundation models have transformed artificial intelligence by enabling large-scale pretraining on diverse datasets, which supports a wide range of subsequent tasks. While many of the most capable models remain closed-access, the emergency of open-weight foundation model where trained parameters are publicly released has facilitated research reproducibility, domain adaptation and community driven innovation. In this survey, we provide a systematic overview of open-weight models across language, vision and multimodal modalities. We introduce a taxonomy based on architecture, modality and openness, examine transformer-based foundations and modern training pipelines and provide a comparative analysis of major model families including LLaMA, Mistral, Falcon, BLOOM, Stable Diffusion and LLaVA-style architectures. We also review standard and multimodal evaluation benchmarks, discuss performance gaps versus closed models and highlight emerging challenges related to security, misuse, compute inequality and AI sovereignty. By synthesizing current developments and technical foundations, this survey offers researchers, practitioners and policymakers a structured framework for understanding the capabilities, limitations and strategic implications of open-weight foundation models.
Paper
Full text
Exploring Open-Weight Foundation Models: A Systematic Review from Transformers to Multimodal AI
Semantic Scholar · Computer Science · 2026
Abstract
Foundation models have transformed artificial intelligence by enabling large-scale pretraining on diverse datasets, which supports a wide range of subsequent tasks. While many of the most capable models remain closed-access, the emergency of open-weight foundation model where trained parameters are publicly released has facilitated research reproducibility, domain adaptation and community driven innovation. In this survey, we provide a systematic overview of open-weight models across language, vision and multimodal modalities. We introduce a taxonomy based on architecture, modality and openness, examine transformer-based foundations and modern training pipelines and provide a comparative analysis of major model families including LLaMA, Mistral, Falcon, BLOOM, Stable Diffusion and LLaVA-style architectures. We also review standard and multimodal evaluation benchmarks, discuss performance gaps versus closed models and highlight emerging challenges related to security, misuse, compute inequality and AI sovereignty. By synthesizing current developments and technical foundations, this survey offers researchers, practitioners and policymakers a structured framework for understanding the capabilities, limitations and strategic implications of open-weight foundation models.