Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation
Perceptual ambiguity and task conflict limit multi-task robotic manipulation via imitation learning. We propose a framework combining a Language-Conditioned Visual Representation (LCVR) module and a Language-conditioned Mixture-of-Experts Density Policy (LMoE-DP). LCVR resolves perceptual ambiguity by grounding visual features with language instructions, enabling differentiation between visually similar tasks. To mitigate task conflict, LMoE-DP uses a sparse expert architecture to specialize in distinct, multimodal action distributions, stabilized by gradient modulation. Experiments in simulation and on a real robot show consistent improvements over strong multi-task baselines. LMoE-DP achieves 73.4% success on LIBERO and 79% on real-robot tasks. Ablations isolate the contributions of LCVR, sparse specialization, and gradient modulation. Together, these components enable efficient and robust multi-task manipulation.
Paper
References (48)
Scroll for more · 36 remaining