Performance of LLMs on Stochastic Modeling Operations Research Problems: From Theory to Practice
Large language models (LLMs) have exhibited capabilities comparable to those of human experts in various fields. However, their modeling abilities-the process of converting real-life problems (or their verbal descriptions) into sensible mathematical models-have been underexplored. In this work, we take the first step to evaluate LLMs' abilities to solve stochastic modeling problems, a model class at the core of Operations Research (OR) and decision-making more broadly. We manually procure a representative set of graduate-level homework and doctoral qualification-exam problems and test LLMs' abilities to solve them. We further leverage SimOpt, an open-source library of simulation-optimization problems and solvers, to investigate LLMs' abilities to make real-world decisions. Our results show that, though a nontrivial amount of work is still needed to reliably automate the stochastic modeling pipeline in reality, state-of-the-art LLMs demonstrate proficiency on par with human experts in both classroom and practical settings.