We propose to leverage Transformer architectures for non-autoregressive human\nmotion prediction. Our approach decodes elements in parallel from a query\nsequence, instead of conditioning on previous predictions such as\ninstate-of-the-art RNN-based approaches. In such a way our approach is less\ncomputational intensive and potentially avoids error accumulation to long term\nelements in the sequence. In that context, our contributions are fourfold: (i)\nwe frame human motion prediction as a sequence-to-sequence problem and propose\na non-autoregressive Transformer to infer the sequences of poses in parallel;\n(ii) we propose to decode sequences of 3D poses from a query sequence generated\nin advance with elements from the input sequence;(iii) we propose to perform\nskeleton-based activity classification from the encoder memory, in the hope\nthat identifying the activity can improve predictions;(iv) we show that despite\nits simplicity, our approach achieves competitive results in two public\ndatasets, although surprisingly more for short term predictions rather than for\nlong term ones.\n
Paper
References (24)
Scroll for more · 12 remaining