Summary
New benchmark for identifying cross cultural similarity of traditional concepts. Through semi automatic annotation of fine grained features for food and clothing concepts, concepts similarity is defined by shared features. LLMs are evaluated on distinguishing cross cultural concepts with high similarity versus low, a novel task as a step towards cross cultural communication facilitated by LLMs.
Reasons to accept
Well motivated, interesting and even groundbreaking goal statement, to enable cross cultural understanding and common ground facilitated by LLMs. The proposed task effectively operationalizes one aspect of this seemingly intractable goal.
Expansion of traditional cross lingual semantic matching (lexicon induction etc) to the pragmatic level by focusing on concepts used in similar contexts.
Creative and effective annotation process, where fine grained features are used to quantify concepts similarity for relatively objective quantification. There is very high agreement for the features annotation.
Interesting results showing that LLMs lag behind humans in the purposes task.
Reasons to reject
Part of the annotation process is automated with ChatGPT, but the version is not specified, and neither is any indication of the performance (quality) of that automated step. Namely, it is used to summarize Wikipedia articles to extract excerpts relevant to the defined features; but we don't know what is the recall compared to a human doing this task or if there's a performance bias towards concepts from certain countries.
While the idea of contrasive classification by number of shared features is interesting, I am not convinced it is justified to treat it as ground truth. Some concepts may be perceived by a well versed human as very similar despite having only a few shared features, maybe because some features are more important than others. And vice versa. The study here does not corroborate the calculated similarity with human judgments.
Unclear experimental setup: a setting is mentioned where concepts names are masked and only features are available, but no results for this setting are presented in the main paper. Further, it seems odd to use this setup, since it seems like it relies purely on the arithmetic ability to calculate the Jaccard index from the lists of features, and does not require any cultural awareness. This also applies to the setting where both features and concepts are provided, since high performance can simply be explained by mathematical ability.
No details about the annotators except that they are "from Eastern countries".
The feature lists seem somewhat arbitrary and are not well motivated. For example many Jewish holidays are features, though they're not necessarily very cross cultural.
Questions to authors
"LLMs are still limited to capturing" should be "LLMs are still limited in their ability to capture"?
"construct this feature list at the same while annotating" - redundant "at the same time"?
In the framework of Hershcovich et al. [1], the contribution here is limited to the Common Ground dimension. What about the other dimensions? Would facilitating cross cultural communication require addressing Aboutness and Values too?
Is each matrix in figure 3 symmetric? It may be easier to read if you drop everything under the diagonal to avoid repetition.
Did you try Llama 3? It may be worth adding it to the next version. Also GPT-4.
In figure 4, "significantly" refers to what p value and what statistical test?
[1] Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders Søgaard. 2022. Challenges and Strategies in Cross-Cultural NLP. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6997–7013, Dublin, Ireland. Association for Computational Linguistics.
Ethics concerns details
No information on annotators or compensation