A key challenge in robotic food manipulation is modeling the material\nproperties of diverse and deformable food items. We propose using a multimodal\nsensory approach to interact and play with food that facilitates the ability to\ndistinguish these properties across food items. First, we use a robotic arm and\nan array of sensors, which are synchronized using ROS, to collect a diverse\ndataset consisting of 21 unique food items with varying slices and properties.\nAfterwards, we learn visual embedding networks that utilize a combination of\nproprioceptive, audio, and visual data to encode similarities among food items\nusing a triplet loss formulation. Our evaluations show that embeddings learned\nthrough interactions can successfully increase performance in a wide range of\nmaterial and shape classification tasks. We envision that these learned\nembeddings can be utilized as a basis for planning and selecting optimal\nparameters for more material-aware robotic food manipulation skills.\nFurthermore, we hope to stimulate further innovations in the field of food\nrobotics by sharing this food playing dataset with the research community.\n