Exploring PCA-based feature representations of image pixels via CNN to enhance food image segmentation

For open-vocabulary recognition of ingredients in food images, ingredient segmentation is a crucial step; however, it remains a challenging problem. This paper proposes a novel approach that explores latent feature representations of image pixels via a convolutional neural network (CNN) to enhance ingredient segmentation. An internal clustering metric based on the silhouette score is defined to evaluate the clustering quality of various pixel-level feature representations generated by different feature maps derived from various CNN backbones. Using this metric, the paper explores the most suitable feature representations and clustering methods for ingredient segmentation. Additionally, it is found that principle component (PC)-driven latent feature presentations of pixels derived from concatenations of backbone feature maps improve the clustering quality of them, resulting in stable segmentation outcomes. Notably, the number of selected eigenvalues can be used as the number of clusters to achieve good segmentation results. The proposed method is validated on widely used public datasets, achieving state-of-the-art performances. Importantly, the proposed segmentation method is unsupervised, and pixel-level feature representations from backbones are not fine-tuned on specific datasets. This demonstrates the flexibility, generalizability, and interpretability of the proposed method, while reducing the need for extensive labeled datasets.

Paper

References (26)

Scroll for more · 14 remaining

Similar papers

© 2026 NYSGPT2525 LLC