Supervised Training of Dense Object Nets using Optimal Descriptors for Industrial Robotic Applications

Dense Object Nets (DONs) by Florence, Manuelli and Tedrake (2018) introduced\ndense object descriptors as a novel visual object representation for the\nrobotics community. It is suitable for many applications including object\ngrasping, policy learning, etc. DONs map an RGB image depicting an object into\na descriptor space image, which implicitly encodes key features of an object\ninvariant to the relative camera pose. Impressively, the self-supervised\ntraining of DONs can be applied to arbitrary objects and can be evaluated and\ndeployed within hours. However, the training approach relies on accurate depth\nimages and faces challenges with small, reflective objects, typical for\nindustrial settings, when using consumer grade depth cameras. In this paper we\nshow that given a 3D model of an object, we can generate its descriptor space\nimage, which allows for supervised training of DONs. We rely on Laplacian\nEigenmaps (LE) to embed the 3D model of an object into an optimally generated\nspace. While our approach uses more domain knowledge, it can be efficiently\napplied even for smaller and reflective objects, as it does not rely on depth\ninformation. We compare the training methods on generating 6D grasps for\nindustrial objects and show that our novel supervised training approach\nimproves the pick-and-place performance in industry-relevant tasks.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC