Attention-Based Fusion Factor to FPN for Small Object Detection

FPN-based detectors have made notable progress in general object detection, but the performance and efficiency of detecting small objects are far from satisfactory. Even through FPN can fuse semantic information from deep layers to shallow layers, it also brings noise containing spatial information of large objects to shallow layers. To address this challenge, we propose an attention-based fusion factor (AFF) consisting of two components, <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\boldsymbol{\alpha}-\mathbf{CBAM}$</tex> and <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\boldsymbol{\alpha}-\mathbf{Query}$</tex> , to enhance the semantic information of small objects in the fusion process. The <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\boldsymbol{\alpha}-\mathbf{CBAM}$</tex> can be used to learn the vertical fusion factor of the FPN to control the information passed from deep to shallow layers. The <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\boldsymbol{\alpha}-\mathbf{Query}$</tex> acts as an fusion factor to control the horizontal transfer impact of FPN. <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\boldsymbol{\alpha}-\mathbf{Query}$</tex> first predicts the coarse locations of small objects in low-resolution feature maps and then increases their weights in high-resolution feature maps. Ablation studies on the publicly available dataset PASCAL VOC are conducted to demonstrate the effectiveness of <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\boldsymbol{\alpha}-\mathbf{CBAM}$</tex> and <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\boldsymbol{\alpha}-\mathbf{Query}$</tex> . Moreover, experimental results on the datasets MS COCO and VisDrone show that the baseline model with additional AFF evidently outperforms itself in almost all the related metrics.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC