Natural Selection via Foundation Models for Soft Robot Evolution

Designing soft robots is a complex and iterative process that demands cross-disciplinary expertise in materials science, mechanics, and control. While foundation models, especially large language models (LLMs), have demonstrated strong reasoning ability, their capability for embodied design evaluation remains underexplored. We introduce RoboCrafter-QA, a benchmark that evaluates whether LLMs can map high-level task descriptions to low-level morphological and material choices by selecting the better design from paired candidates in EvoGym. The benchmark spans 12 locomotion, manipulation, and balancing tasks, and explicitly controls difficulty by the reward gap between candidate designs. Evaluating multiple frontier LLMs, we find that they can distinguish designs with large performance gaps but struggle with fine-grained comparisons when rewards are similar. To address this, we finetune a compact open-source LLM using parameter-efficient adaptation and achieve strong performance on RoboCrafter-QA. We further show that the finetuned model can be prompted to generate feasible, high-performing morphologies under actuator constraints. Finally, we fabricate a modular soft robot platform and empirically validate that designs with higher simulated reward tend to yield higher real-world speed, supporting the practical relevance of benchmark performance. Our code and data are available at https://github.com/robocrafterqa/robocrafterqa_code.

Paper

Similar papers

© 2026 NYSGPT2525 LLC