Pedestrian detection has been considered essential for many real-world applications. Continuous progress in the detection accuracy tends to be supported by solutions of increasing complexity, which inflicts scaling processing costs. In this paper, we counter that tendency by exploring the fastest generic object detector, YOLOv3, by applying in the learning phase the weak semantic segmentation infusion technique - inspired in the SD- SRCNN - increasing its accuracy without affecting the inference speed. Experiments are run on the Caltech Pedestrian Detection Benchmark, as well as in a real-world video surveillance scenario with the PTI01 Pedestrian Detection Dataset. Our proposed method demonstrates infusion benefits in some scenarios achieving competitive accuracy on Caltech, staying less than 3% behind the best method on the primary metric, while being 5 faster. We verify that, among the compared methods, YOLOv3 seems to be the most feasible one for practical applications. Lastly, we also contribute by providing improved annotations for the PTI01 dataset.
Paper
Full text
Detecting Pedestrians with YOLOv3 and Semantic Segmentation Infusion
Semantic Scholar · Computer Science · 2019
Abstract
Pedestrian detection has been considered essential for many real-world applications. Continuous progress in the detection accuracy tends to be supported by solutions of increasing complexity, which inflicts scaling processing costs. In this paper, we counter that tendency by exploring the fastest generic object detector, YOLOv3, by applying in the learning phase the weak semantic segmentation infusion technique - inspired in the SD- SRCNN - increasing its accuracy without affecting the inference speed. Experiments are run on the Caltech Pedestrian Detection Benchmark, as well as in a real-world video surveillance scenario with the PTI01 Pedestrian Detection Dataset. Our proposed method demonstrates infusion benefits in some scenarios achieving competitive accuracy on Caltech, staying less than 3% behind the best method on the primary metric, while being 5 faster. We verify that, among the compared methods, YOLOv3 seems to be the most feasible one for practical applications. Lastly, we also contribute by providing improved annotations for the PTI01 dataset.