Can LLMs Solve Solubility Tasks? The SoluBench Benchmark for Pure and Mixed Solvent Systems.

Solubility prediction is a fundamental task in chemistry and drug discovery, yet the capability of large language models (LLMs) to solve solubility-related problems remains underexplored. Here, we present SoluBench, a benchmark comprising 9806 questions across four tasks of increasing complexity, designed to evaluate LLMs' performance on solubility tasks in both pure and mixed solvent systems. The tasks include pairwise solvent comparison (Task 1), single-best solvent selection (Task 2), cosolvent effect prediction (Task 3), and pairwise compound comparison (Task 4). The benchmark is grounded in experimental data from BigSolDB 2.0 and MixtureSolDB data sets. Systematic evaluation of over 20 proprietary and open-source LLMs reveals that frontier proprietary models achieve strong performance on solvent-oriented tasks, with Gemini 3 Flash reaching 90.6%, 66.2%, and 86.4% accuracy on Tasks 1, 2, and 3, respectively. The solute-focused Task 4 is exceptionally challenging in standard settings, requiring explicit reasoning to achieve functional accuracy, particularly for polar aprotic solvents such as DMSO and DMF. Across all tasks, a universal performance bottleneck is identified for large solutes (MolWt > 500 Da). These findings demonstrate that while LLMs possess intrinsic chemical analysis capabilities sufficient for qualitative solubility assessment in technological practice, significant limitations in molecular structure understanding persist.

Paper

Full text

PDF

Can LLMs Solve Solubility Tasks? The SoluBench Benchmark for Pure and Mixed Solvent Systems.

Semantic Scholar · Chemistry · 2026

Abstract

Solubility prediction is a fundamental task in chemistry and drug discovery, yet the capability of large language models (LLMs) to solve solubility-related problems remains underexplored. Here, we present SoluBench, a benchmark comprising 9806 questions across four tasks of increasing complexity, designed to evaluate LLMs' performance on solubility tasks in both pure and mixed solvent systems. The tasks include pairwise solvent comparison (Task 1), single-best solvent selection (Task 2), cosolvent effect prediction (Task 3), and pairwise compound comparison (Task 4). The benchmark is grounded in experimental data from BigSolDB 2.0 and MixtureSolDB data sets. Systematic evaluation of over 20 proprietary and open-source LLMs reveals that frontier proprietary models achieve strong performance on solvent-oriented tasks, with Gemini 3 Flash reaching 90.6%, 66.2%, and 86.4% accuracy on Tasks 1, 2, and 3, respectively. The solute-focused Task 4 is exceptionally challenging in standard settings, requiring explicit reasoning to achieve functional accuracy, particularly for polar aprotic solvents such as DMSO and DMF. Across all tasks, a universal performance bottleneck is identified for large solutes (MolWt > 500 Da). These findings demonstrate that while LLMs possess intrinsic chemical analysis capabilities sufficient for qualitative solubility assessment in technological practice, significant limitations in molecular structure understanding persist.

Similar papers

© 2026 NYSGPT2525 LLC