Although there are several food-level benchmarks for food-related learning, the lack of fine-grained ingredient annotation significantly impedes progress in food scene understanding. In this study, we focus on Chinese food understanding which involves fine-grained ingredient detection and cross-modal ingredient retrieval. Specifically, to support studies on Chinese food understanding, we build the first cross-modal ingredient-level dataset called CMIngre, which contains 8,001 image-text pairs from three different sources, i.e. dishes, recipes, and user-generated content, covering 429 distinct ingredients and 95,290 bounding boxes. Based on CMIngre, we evaluate the performance of traditional CNN-based detection algorithms and transformer-based pre-trained large models for ingredient detection. We also propose baseline methods for the cross-modal ingredient retrieval task in both the end-to-end and two-stage settings. Extensive experiments on CMIngre demonstrate the effectiveness of our proposed methods on food understanding.
Paper
Full text
Toward Chinese Food Understanding: A Cross-Modal Ingredient-Level Benchmark
Semantic Scholar · Agricultural and Food Sciences · 2025
Abstract
Although there are several food-level benchmarks for food-related learning, the lack of fine-grained ingredient annotation significantly impedes progress in food scene understanding. In this study, we focus on Chinese food understanding which involves fine-grained ingredient detection and cross-modal ingredient retrieval. Specifically, to support studies on Chinese food understanding, we build the first cross-modal ingredient-level dataset called CMIngre, which contains 8,001 image-text pairs from three different sources, i.e. dishes, recipes, and user-generated content, covering 429 distinct ingredients and 95,290 bounding boxes. Based on CMIngre, we evaluate the performance of traditional CNN-based detection algorithms and transformer-based pre-trained large models for ingredient detection. We also propose baseline methods for the cross-modal ingredient retrieval task in both the end-to-end and two-stage settings. Extensive experiments on CMIngre demonstrate the effectiveness of our proposed methods on food understanding.