Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation

Perceptual ambiguity and task conflict limit multi-task robotic manipulation via imitation learning. We propose a framework combining a Language-Conditioned Visual Representation (LCVR) module and a Language-conditioned Mixture-of-Experts Density Policy (LMoE-DP). LCVR resolves perceptual ambiguity by grounding visual features with language instructions, enabling differentiation between visually similar tasks. To mitigate task conflict, LMoE-DP uses a sparse expert architecture to specialize in distinct, multimodal action distributions, stabilized by gradient modulation. Experiments in simulation and on a real robot show consistent improvements over strong multi-task baselines. LMoE-DP achieves 73.4% success on LIBERO and 79% on real-robot tasks. Ablations isolate the contributions of LCVR, sparse specialization, and gradient modulation. Together, these components enable efficient and robust multi-task manipulation.

Paper

References (48)

Scroll for more · 36 remaining

Similar papers

© 2026 NYSGPT2525 LLC