Anime Character Style Classification Based on Frequency-Domain Decoupling and Multi-Scale Feature Fusion

Automatic classification of anime character painting styles is of great significance to the digital cultural industry and visual content production. Existing methods are prone to shortcut learning when handling complex color rendering and cannot fully decouple high-frequency line drafts from low-frequency colors. To solve this problem, this study proposes an improved deep learning classification method based on EfficientNetV2-B0. This method introduces random amplitude scaling (RAS) at the data input terminal. It realizes effective decoupling of colors and line-draft structures through random low-frequency amplitude perturbation, and suppresses the model’s excessive dependence on global color information from the source. Edge-guided coordinate attention (EG-CA) is integrated into the backbone network. It enhances the perception of line and contour features through edge weights and improves the model’s ability to capture fine-grained structural features. Adaptive scale feature aggregation (ASFA) is designed in the multi-scale feature fusion stage. It achieves efficient fusion of shallow textures and deep semantics through dynamic weighting, so as to enhance the model’s discriminative ability under complex painting styles. On a dataset containing 7887 images of four categories, the classification accuracy of the model reaches 95.81%. It significantly outperforms mainstream models such as MViTv2-T. Meanwhile, the number of parameters is only 7.84 M and the inference speed reaches 68.83 FPS. Ablation experiments show that the synergistic effect of the three modules improves the accuracy of the baseline model by 6.06%. It proves that the proposed method provides reliable technical support for the structured management and copyright traceability of anime images.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC