Response to Weaknesses-2&3
**Response to Weaknesses-2:**
Thanks for your comments.
As suggested, we introduce the Periodically Exchange Teacher-Student (PETS) [Ref 1] and Target Prediction Distribution Searching (TPDS) [Ref 2] as the comparison methods.
As the table below shows, KGD consistently outperforms PETS [Ref 1] and TPDS [Ref 2].
In the tables below, $\star$ signifies that the methods employ WordNet to retrieve category definitions given category names, and CLIP to predict classification pseudo labels for objects. We adopt AP50 in evaluations. The results of all methods are acquired with the same baseline (Detic[81]) as shown in the first row.
We will include the new experiments into the updated paper later. Thank you for your suggestion!
[Ref 1] Liu, Qipeng, et al. "Periodically exchange teacher-student for source-free object detection." Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023.
[Ref 2] Tang, Song, et al. "Source-free domain adaptation via target prediction distribution searching." International journal of computer vision 132.3 (2024): 654-672.
Method |Cityscapes |Vistas |BDD100K-weather-rain|BDD100K-weather-snow|BDD100K-weather-overcast|BDD100K-weather-cloudy|BDD100K-weather-foggy |BDD100K-time-of-day--daytime |BDD100K-time-of-day--dawn&dusk|BDD100K-time-of-day--night
-|-|-|-|-|-|-|-|-|-|-
Detic (Baseline) |46.5 |35.0 |34.3 |33.5 |39.1 |42.0 |28.4 |39.2 |35.3 |28.5
PETS |50.2 |35.8 |34.4 |33.9 |40.1 |43.0 |36.3 |39.7 |35.7 |27.8
PETS$\star$ |50.8 |37.4 |35.9 |36.3 |41.0 |42.8 |36.7 |40.9 |37.2 |27.7
TPDS |50.1 |36.0 |35.8 |35.2 |40.0 |42.1 |36.4 |40.4 |36.5 |28.5
TPDS$\star$ |50.3 |37.1 |35.6 |35.9 |40.5 |43.4 |36.9 |41.3 |36.7 |28.9
KGD (Ours) |53.6 |40.3 |37.3 |37.1 |44.6 |48.2 |38.0 |46.6 |41.0 |31.2
Method |Common Objects: VOC|Common Objects: Objects365 |Intelligent Surveillance: MIO-TCD |Intelligent Surveillance:BAAI |Intelligent Surveillance:VisDrone |Artistic Illustration:Clipart1k|Artistic Illustration:Watercolor2k|Artistic Illustration:Comic2k
-|-|-|-|-|-|-|-|-
Detic (Baseline) |83.9 |29.4 |20.6 |20.6 |19.0 |61.0 |58.9 |51.2
PETS |85.9 |31.5 |20.6 |22.6 |18.2 |63.0 |60.2 |50.4
PETS$\star$ |86.3 |32.1 |21.1 |23.2 |19.3 |63.6 |61.3 |50.6
TPDS |85.5 |31.8 |20.2 |22.1 |18.8 |63.1 |60.0 |50.1
TPDS$\star$ |85.6 |32.0 |21.1 |23.2 |19.2 |64.3 |61.4 |50.6
KGD (Ours) |86.9 |34.4 |24.6 |24.3 |23.7 |69.1 |63.5 |55.6
**Response to Weaknesses-3:**
Thanks for your comment. As suggested, we conduct new experiments to compare our KGD with prior knowledge graph distillation methods[Ref 3, Ref 4] . The results in the tables below show that our KGD outperform [Ref 3, Ref 4] clearly, largely becuase the knowledge graphs in [Ref 3, Ref 4] are hand-crafted by domain experts while ours is built and learnt from CLIP. We will include the new experiments into the updated paper later.
Method |Cityscapes |Vistas |BDD100K-weather-rain|BDD100K-weather-snow|BDD100K-weather-overcast|BDD100K-weather-cloudy|BDD100K-weather-foggy |BDD100K-time-of-day--daytime |BDD100K-time-of-day--dawn&dusk|BDD100K-time-of-day--night
-|-|-|-|-|-|-|-|-|-|-
Detic (Baseline) |46.5 |35.0 |34.3 |33.5 |39.1 |42.0 |28.4 |39.2 |35.3 |28.5
KGE [Ref 5] |48.9 |36.0 |35.5 |34.4 |40.5 |41.2 |29.7 |40.1 |36.6 |29.0
Context Matters [Ref 6] |49.4 |36.6 |36.3 |35.0 |41.7 |42.4 |30.2 |41.5 |37.2 |29.7
KGD (Ours) |53.6 |40.3 |37.3 |37.1 |44.6 |48.2 |38.0 |46.6 |41.0 |31.2
Method |Common Objects: VOC|Common Objects: Objects365 |Intelligent Surveillance: MIO-TCD |Intelligent Surveillance:BAAI |Intelligent Surveillance:VisDrone |Artistic Illustration:Clipart1k|Artistic Illustration:Watercolor2k|Artistic Illustration:Comic2k
-|-|-|-|-|-|-|-|-
Detic (Baseline) |83.9 |29.4 |20.6 |20.6 |19.0 |61.0 |58.9 |51.2
KGE [Ref 5] |85.4 |31.2 |20.3 |23.5 |19.4 |62.4 |58.1 |50.5
Context Matters [Ref 6] |85.9 |31.7 |20.9 |23.3 |19.9 |62.9 |59.1 |52.3
KGD (Ours) |86.9 |34.4 |24.6 |24.3 |23.7 |69.1 |63.5 |55.6
[Ref 3] Christopher Lang, Alexander Braun, and Abhinav Valada. Contrastive object detection using knowledge graph embeddings. arXiv preprint arXiv:2112.11366, 2021
[Ref 4] Aijia Yang, Sihao Lin, Chung-Hsing Yeh, Minglei Shu, Yi Yang, and Xiaojun Chang. Context matters: Distilling knowledge graph for enhanced object detection. IEEE Transactions on Multimedia, 2023, doi: 10.1109/TMM.2023.3266897.