CAME: Convolution and Attention Construct Multi-Scale Neural Network Efficiently for Medical Image Classification

The study of medical image classification is of great importance to assist doctors in diagnosing conditions and intelligently identifying types of illnesses. However, unlike general image classification, this task is still very challenging because medical images have more complex and variable structural patterns. There are two directions currently explored: transfer learning and Transformer-based, but they have certain drawbacks. For example, transfer learning requires a large amount of annotated data support and the training process is tedious and complicated; the Transformer-based approach has the problem of high computational time complexity, which is also very time-consuming. It cannot be ignored that their representations for classification need to be enhanced. Based on the above issues, our proposed CAME consists of three feature extraction modules from different scales including local feature information, global semantic representation, and external attention, and a multi-scale feature aggregation module (MSFA). The MSFA module enhances the semantics of each of the three scales of representations through space, through the attention mechanism, and then aggregates the three enhanced representations to obtain the final representation. In the experiments, the proposed CAME performs best in the baseline on the benchmarks of the three criteria and achieves end-to-end medical image classification.

Paper

Full text

PDF

CAME: Convolution and Attention Construct Multi-Scale Neural Network Efficiently for Medical Image Classification

Semantic Scholar · Computer Science · 2023

Abstract

The study of medical image classification is of great importance to assist doctors in diagnosing conditions and intelligently identifying types of illnesses. However, unlike general image classification, this task is still very challenging because medical images have more complex and variable structural patterns. There are two directions currently explored: transfer learning and Transformer-based, but they have certain drawbacks. For example, transfer learning requires a large amount of annotated data support and the training process is tedious and complicated; the Transformer-based approach has the problem of high computational time complexity, which is also very time-consuming. It cannot be ignored that their representations for classification need to be enhanced. Based on the above issues, our proposed CAME consists of three feature extraction modules from different scales including local feature information, global semantic representation, and external attention, and a multi-scale feature aggregation module (MSFA). The MSFA module enhances the semantics of each of the three scales of representations through space, through the attention mechanism, and then aggregates the three enhanced representations to obtain the final representation. In the experiments, the proposed CAME performs best in the baseline on the benchmarks of the three criteria and achieves end-to-end medical image classification.

Similar papers

© 2026 NYSGPT2525 LLC