GIFT: Generalizable Interaction-aware Functional Tool Affordances without Labels

Tool use requires reasoning about the fit between an object's affordances and\nthe demands of a task. Visual affordance learning can benefit from\ngoal-directed interaction experience, but current techniques rely on human\nlabels or expert demonstrations to generate this data. In this paper, we\ndescribe a method that grounds affordances in physical interactions instead,\nthus removing the need for human labels or expert policies. We use an efficient\nsampling-based method to generate successful trajectories that provide contact\ndata, which are then used to reveal affordance representations. Our framework,\nGIFT, operates in two phases: first, we discover visual affordances from\ngoal-directed interaction with a set of procedurally generated tools; second,\nwe train a model to predict new instances of the discovered affordances on\nnovel tools in a self-supervised fashion. In our experiments, we show that GIFT\ncan leverage a sparse keypoint representation to predict grasp and interaction\npoints to accommodate multiple tasks, such as hooking, reaching, and hammering.\nGIFT outperforms baselines on all tasks and matches a human oracle on two of\nthree tasks using novel tools.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC