Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control

Deep neural network (DNN)-based policy models, such as vision-language-action (VLA) models, excel at automating complex decision-making from multi-modal inputs. However, scaling these models greatly increases computational overhead, complicating deployment in resourceconstrained settings like robot manipulation and autonomous driving. To address this, we propose SaliencyAware Quantized Imitation Learning (SQIL), which combines quantization-aware training with a selective lossweighting strategy for mission-critical states. By identifying these states via saliency scores and emphasizing them in the training loss, SQIL preserves decision fidelity under low-bit precision. We validate SQIL's generalization capability across extensive simulation benchmarks with environment variations, real-world tasks, and cross-domain tasks (self-driving, physics simulation), consistently recovering full-precision performance. Notably, a 4-bit weightquantized VLA model for robotic manipulation achieves up to $2.5 \times$ speedup and $2.5 \times$ energy savings on an edge GPU with minimal accuracy loss. These results underline SQIL 's potential for efficiently deploying large IL-based policy models on resource-limited devices.

Paper

References (60)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC