Technical Report: Disentangled Action Parsing Networks for Accurate Part-level Action Parsing
Part-level Action Parsing aims at part state parsing for boosting action\nrecognition in videos. Despite of dramatic progresses in the area of video\nclassification research, a severe problem faced by the community is that the\ndetailed understanding of human actions is ignored. Our motivation is that\nparsing human actions needs to build models that focus on the specific problem.\nWe present a simple yet effective approach, named disentangled action parsing\n(DAP). Specifically, we divided the part-level action parsing into three\nstages: 1) person detection, where a person detector is adopted to detect all\npersons from videos as well as performs instance-level action recognition; 2)\nPart parsing, where a part-parsing model is proposed to recognize human parts\nfrom detected person images; and 3) Action parsing, where a multi-modal action\nparsing network is used to parse action category conditioning on all detection\nresults that are obtained from previous stages. With these three major models\napplied, our approach of DAP records a global mean of $0.605$ score in 2021\nKinetics-TPS Challenge.\n