Exploiting latent representation of sparse semantic layers for improved short-term motion prediction with Capsule Networks
As urban environments manifest high levels of complexity it is of vital\nimportance that safety systems embedded within autonomous vehicles (AVs) are\nable to accurately anticipate short-term future motion of nearby agents. This\nproblem can be further understood as generating a sequence of coordinates\ndescribing the future motion of the tracked agent. Various proposed approaches\ndemonstrate significant benefits of using a rasterised top-down image of the\nroad, with a combination of Convolutional Neural Networks (CNNs), for\nextraction of relevant features that define the road structure (eg. driveable\nareas, lanes, walkways). In contrast, this paper explores use of Capsule\nNetworks (CapsNets) in the context of learning a hierarchical representation of\nsparse semantic layers corresponding to small regions of the High-Definition\n(HD) map. Each region of the map is dismantled into separate geometrical layers\nthat are extracted with respect to the agent's current position. By using an\narchitecture based on CapsNets the model is able to retain hierarchical\nrelationships between detected features within images whilst also preventing\nloss of spatial data often caused by the pooling operation. We train and\nevaluate our model on publicly available dataset nuTonomy scenes and compare it\nto recently published methods. We show that our model achieves significant\nimprovement over recently published works on deterministic prediction, whilst\ndrastically reducing the overall size of the network.\n
Paper
References (33)
Scroll for more · 21 remaining