Drawing on 1,178 AI risk and reliability papers from 9,439 generative AI papers (January 2020 through March 2025), we compare research outputs from the most influential research institutions: frontier AI companies (Anthropic, Google DeepMind, Meta, Microsoft, and OpenAI) and leading AI universities (CMU, MIT, NYU, Stanford, UC Berkeley, and University of Washington). We find that Frontier Corporate AI research increasingly concentrates on pre-deployment areas — model alignment and testing & evaluation — while attention to deployment-stage issues, such as model bias, has waned as commercial imperatives and existential risk concerns have taken precedence. We identify significant research gaps in high-risk deployment domains, including healthcare applications, commercial and financial contexts, misinformation, persuasive and addictive features, hallucinations, and copyright usage in training and inference. Without concerted efforts to enhance external observability into AI’s deployment, the growing concentration of AI research with frontier corporations could deepen knowledge deficits in these critical deployment areas. We recommend measures to expand external researcher access to deployment data and improve systematic observability of AI systems’ in-market behaviors.
Paper
References (100)
Scroll for more · 38 remaining