Summary
This work addresses the issue of estimating the importance of on-road objects using video sequences from a driver’s perspective, a critical task for enhancing driving safety. The authors introduce the Traffic Object Importance (TOI) dataset, which is significantly larger and more diverse than existing datasets, and propose a novel model that integrates multi-fold top-down guidance factors—driver intention, semantic context, and traffic rules—with bottom-up features for more accurate importance estimation. Experimental results demonstrate that the proposed model significantly outperforms state-of-the-art methods in on-road object importance estimation.
Strengths
1. The introduction of the Traffic Object Importance (TOI) dataset, which is significantly larger and more diverse than existing datasets, provides a robust foundation for training and evaluating models in on-road object importance estimation, thereby addressing a major limitation in the field.
2. The proposed model effectively integrates multi-fold top-down guidance factors—driver intention, semantic context, and traffic rules—with bottom-up features, which showed good performance for the TOI task.
Weaknesses
1. Lack of description of the annotation details.
How many annotators are involved in the annotation procedure? It would be good if the authors can provide some annotation procedure samples regarding the double-checking annotation mechanism and the triple-discussing annotation mechanism.
2. It seems this annotation will be varied according to different traffic rules. Since KITTI is collected in Germany, the annotators should be familiar to germany traffic rules. However the authors did not mention this information in their submission, thereby the label quality is doubtful.
3. The authors are encouraged to build up the first benchmark based on the proposed dataset by using various existing object detection methods, e.g., Yolo, with the proposed head or simpler head. It is interesting to see how the existing object detectors work on this new task.
4. More statistics of the dataset are encouraged to be given, e.g., the number of important object of different categories, etc.
Questions
1. How many annotators were involved in the annotation procedure for the dataset? Can the authors provide detailed examples of their double-checking and triple-discussing annotation mechanisms?
2. Were the annotators familiar with German traffic rules, given that the dataset was collected in Germany (KITTI dataset)? How was the expertise of the annotators in relation to German traffic laws ensured and validated?
3. Have the authors considered building the first benchmark using their dataset with existing object detection methods, such as YOLO? What were the performance outcomes of these existing methods when applied to the new task?
4. Can the authors provide more detailed statistics about the dataset, such as the number of important objects in different categories? How do these statistics compare to other datasets in the same domain?
Limitations
yes the authors mentioned it in the appendix