Measuring Taiwanese Mandarin Language Understanding

The evaluation of large language models (LLMs) has drawn substantial attention in the field recently. This work focuses on evaluating LLMs in a Chinese context, specifically, for Traditional Chinese which has been largely underrepresented in existing benchmarks. We present TMLU, a holistic evaluation suit tailored for assessing the advanced knowledge and reasoning capability in LLMs, under the context of Taiwanese Mandarin. TMLU consists of an array of 37 subjects across social science, STEM, humanities, Taiwan-specific content, and others, ranging from middle school to professional levels. In addition, we curate chain-of-thought-like few-shot explanations for each subject to facilitate the evaluation of complex reasoning skills. To establish a comprehensive baseline, we conduct extensive experiments and analysis on 24 advanced LLMs. The results suggest that Chinese open-weight models demonstrate inferior performance comparing to multilingual proprietary ones, and open-weight models tailored for Taiwanese Mandarin lag behind the Simplified-Chinese counterparts. The findings indicate great headrooms for improvement, and emphasize the goal of TMLU to foster the development of localized Taiwanese-Mandarin LLMs. We release the benchmark and evaluation scripts for the community to promote future research.

Paper

Similar papers

Reviewer u2Ve7/10 · confidence 3/52024-05-05

Summary

The paper presents a new evaluation dataset (TMLU) for traditional Mandarin Chinese (Taiwanese). It contains 37 subjects (social science, STEM, humanities and Taiwan specific). The contributions of this eval dataset and this paper are: It aims at better evaluating LLMs on traditional Chinese, which has unique characteristics and challenges compared to simplified Chinese Compared to existing traditional Chinese evaluation datasets (TC-eval and TMMLU-plus), TMLU has the following advantages: 1. not from a single website, using pdfs or words, not easily scrapable so avoiding **data contaminations**; 2. more answerable from the question context; 3. including CoT instructions and Latex expressions

Rating

7

Confidence

3

Ethics flag

1

Reasons to accept

1. It is great to see a new dataset introduced to address the evaluation needs of traditional Chinese (Taiwan) languages. 2. Authors demonstrate TMLU has advantages as an evaluation dataset compared to existing traditional Chinese eval benchmarks, including better handling of data contamination, less unanswerable questions.

Reasons to reject

Although authors show the dataset has advantages compared to existing benchmarks, I hope the dataset can do even better. In particular, I hope the eval dataset can cover the following capabilities: 1. Chatability: being able to understand a dialogue context 2. Understanding nuances from languages 3. Other use cases in addition to exams Please see my questions below.

Questions to authors

1. From my understanding “同志” can mean “comrade” and “homosexual” depending on the context even in simplified Chinese. I still agree with the point that words carry different meanings in simplified and traditional Chinese, and the challenge do exist for LLM evaluations, but this may not be the best example to showcase in the Introduction section? 2. If I understand correctly, the data comes exclusively from examinations (high school, college or professional), that may impose some bias into this dataset in a way that we are encouraging LLMs to be good at answering those “non-daily” questions, but does not necessarily make models more “usable” at answering “daily” questions. So it mainly measures knowledge, not necessarily other aspects of language understanding, such as understanding the nuances in languages and LLMs’ usability and “chattability”. Think about people’s daily usage of LLMs, they may ask questions about trip planning, shopping ideas, financial knowledge, legal advice, emotional support, relationship questions, career advice, business knowledge, etc. I would like to see those covered as well. This could be out of the scope of this paper, though. I hope we will have more such kind of evaluation data in the future. 3. Figure 3 shows that TMLU is less likely to be included in training data of these benchmarked LLMs, but not by a big margin. What are your takeaways on these results?

Reviewer L2nz7/10 · confidence 5/52024-05-10

Summary

This paper describes TMLU, a new MMLU-style dataset for LLM evaluation, targeted specifically at Taiwanese Mandarin (and the traditional Chinese script). As with MMLU and other datasets fashioned around it, the source of the data for TMLU is exam questions from middle and high school, and professional topics (e.g. traditional Chinese medicine or driving exams), with a strong focus on scraping data from PDFs and MSWord documents in the hope that it is less likely to be included in the pre-training data for LLMs. In addition to evaluating a range of open-weight and proprietary LLMs over TMLU, the paper includes some nice analysis of data contamination, direct answering vs. COT prompting, and a breakdown of results across time (based on the timestamps of the documents the questions were sourced from).

Rating

7

Confidence

5

Ethics flag

1

Reasons to accept

- it is important for LLM evaluation resources to be developed for a broad range of languages, and while there are many resources for Mandarin Chinese, there are considerably fewer for Taiwanese Mandarin, and the traditional Chinese script; this combined with the scale of the dataset and attention to detail in the paper makes it worthy of publication - beyond just presenting results for different LLMs over TMLU, the paper includes additional analysis of data contamination (showing that it is lower than a competitor dataset), direct answering vs. COT prompting, and LLM accuracy of questions sourced from different years

Reasons to reject

- not a strong reason to reject, but I would have liked to have seen more comparison of the coverage of topics in the dataset with related datasets such as TMMLU-plus and CMMLU - I would also have liked to have seen more discussion/analysis of results over "Taiwan specific" questions (see below), and also between similar topics in CMMLU and other datasets for Mandarin Chinese - the writing is overall good, but there were some places where things got tangled up, particularly the first sentence of Section 4.1 which should be changed to "We perform comprehensive evaluation over TMLU using 24 different LLMs ..." (you are not evaluating TMLU itself, but rather using TMLU to evaluate LLMs); also "wide range of field*s*" (Section 3.2); "Tradition*al* Chinese (Section 2)

Questions to authors

- overall, models appear to achieve slightly lower results over TMLU than CMMLU, e.g., but how much of this is attributable to script differences (traditional vs. simplified) vs. the Taiwanese/Mandarin Chinese distinction vs. "local knowledge"? For example, if you convert TMLU into the simplified script, how much do the results change? How do you explain the fact that results over "Taiwan specific" questions are actually higher than all other question types? - in what sense is TMLU "holistic" (page 8)? I get that it is broad coverage, but is there a way to quantify the degree of comprehensiveness of the dataset (e.g. relative to MMLU or CMMLU)? - more of a comment than a question, but I found myself getting confused at one point about what TMLU was as compared to TMMLU-plus (i.e. given the similarity between "TMLU" and "TMMLU", I assumed that TMMLU-plus was somehow an extended version of TMLU and tried to go back and find mention of how it had been expanded); perhaps remind the reader in Section 5.1 that they are completely different datasets, despite the similarity in name - what is the basis for you saying that "the applicability of these benchmarks has decreased as the ever larger models are demonstrating human-level performance"? yes, models are improving, but there is a bit of circularity in the slightly awkward wording here, as "human level performance" is being claimed in large part due to improvement over benchmarks, but at the same time, data contamination is becoming a bigger and bigger issue; possibly nuance the claim a bit

Reviewer 8Aze5/10 · confidence 4/52024-05-11

Summary

This paper introduces TMLU, a new evaluation benchmark designed to assess large language models (LLMs) specifically in the context of Taiwanese Mandarin. TMLU addresses the unique linguistic nuances and cultural intricacies of Taiwanese Mandarin, differentiating it from other varieties of Chinese. The benchmark encompasses 37 subjects, spanning middle school to professional levels, and includes approximately 3,000 multiple-choice questions across diverse disciplines. The authors evaluate 24 LLMs on TMLU and include Chain-of-Thought (CoT) examples to assess complex reasoning skills.

Rating

5

Confidence

4

Ethics flag

1

Reasons to accept

The paper introduces a novel benchmark, TMLU, designed specifically for evaluating the performance of large language models (LLMs) in Taiwanese Mandarin. This benchmark encompasses a broad spectrum of subjects and has been meticulously curated through manual review. Additionally, the authors attempt to mitigate data contamination by sourcing questions from PDF and Microsoft Word documents.

Reasons to reject

1. Limited Linguistic and Cultural Differentiation: While the benchmark is written in Traditional Chinese, the questions themselves do not sufficiently capture the unique linguistic and cultural aspects of Taiwanese Mandarin. Many STEM questions, for example, use terminology common to Mainland China. Areas where language usage differences are more pronounced, like social media, are not included. This limitation may explain why Simplified Chinese models perform well on the benchmark. 2. Unclear Discussion and Unconvincing Conclusions: The paper's discussion lacks clarity, and some conclusions are not well-supported. For instance, the authors claim advantages over TMMLU plus, but only provide a few examples without statistical evidence. 3. Limited Innovation: The benchmark construction and analysis are similar to previous work, and the exclusive use of multiple-choice questions limits innovation. The paper would be stronger with a more novel approach to dataset construction and the inclusion of diverse test types. 4. Data Contamination Concerns: While the authors strive to avoid data contamination, their use of materials from major tests in Taiwan raises concerns. These materials may already be available online in digital form on test preparation websites, potentially accessible to LLMs during training. The observed decline in model performance on older questions (pre-2017) could indicate data contamination, as older materials are less likely to have been digitized and disseminated online.

Questions to authors

Figure 2 contains an error in the answer key (the correct answer for the last question is "C"). The example using "Tongzhi" to illustrate language differences is not the most effective, as its primary meaning is consistent across Mainland China and Taiwan (referencing Zhang et al. 2022). Zhang, K., Lu, C., & Zhang, J. (2022). Chinese media representations of tongzhi (2009–2019). Discourse & Communication, 16(3), 305-325.

Reviewer 8Aze2024-06-06

Thank you for pointing out that TMMLU-plus is a concurrent work; I have updated the score accordingly. However, I am still not entirely convinced by the claim that "TMLU is better for the community to objectively evaluate the model capability due to its better transparency, localization, and robustness." While I admit that TMLU has better data quality, as the authors stated, I am not persuaded regarding its localization and robustness. I noticed that TMMLU-plus includes tests on Taiwanese Trivia and Taiwanese Hokkien, which could cover more local language use, particularly informal language.

Authorsrebuttal2024-06-07

Thank you for the insightful comments! Regarding localization — the Taiwanese Trivia question set (TTQA) included in TMMLU-plus is proposed by another prior work [1]. We are also glad to augment TMLU with TTQA in future revision, along with more questions such as Hokkien knowledge to improve the localization of TMLU. Thanks again for your suggestions. [1] Extending the pre-training of bloom for improved support of traditional chinese: Models, Methods and Results Ennen et al. arXiv:2303.04715, 2023

Reviewer 1eqm6/10 · confidence 4/52024-05-14

Summary

This paper proposes a benchmark for traditional Chinese, TMLU, which targets the Taiwan society by construction. This benchmark is like the MMLU format but with traditional Chinese language. Compared with previous benchmarks for Taiwanese Mandarin such as TC-Eval and TMMLU, the authors argue that TMLU differs in that it originates from traditional Chinese and it is sourced from PDF or Word documents that suffer from less contamination. The paper benchmarks many LLMs on the proposed TMLU benchmark.

Rating

6

Confidence

4

Ethics flag

1

Reasons to accept

1. Creating an MMLU-like benchmark for Taiwanese Mandarin is a focused contribution and could benefit the deployment and development of LLMs to serve Taiwan users. 2. Empirical experiments are comprehensive, there are COT and data contamination analysis included besides the main results.

Reasons to reject

1. While the authors try to distinguish TMLU from TMMLU-plus, it is better to have more quantitative evidence given that the two benchmarks are similar. For example, Table 1 notes that the robustness, transparency of TMMLU-plus are low, what do they mean? In Section 2 the authors mention that “In our preliminary investigation, we found that most questions in TMMLU-plus could be found on one single website.”, can you quantify this? How many can be found? 2. In figure 3, I wonder what the absolute values of the y-axis mean, for example, does a min-k% prob of 4.0 indicate that the pretraining data is absolutely likely to present in the pretraining data? Currently the comparison in Figure 3 is relative, and I am not sure whether the gap between TMLU and TMMLU-plus is significant or has meaningful differences given that the gap looks small – maybe TMLU and TMMLU-plus both suffer from data contamination already.

Questions to authors

``` Understanding that TMMLU-Plus is a concurrent work within one month of the submission, I increase my score from 4 to 6 given that I ignore the existence of TMMLU-Plus ```

Reviewer 1eqm2024-06-04

Thank you for the clarifications! Understanding that TMMLU-plus is a concurrent work, I think the contribution of this work is more significant. However, I do suggest that when discussing some drawbacks of TMMLU-plus, it is more convincing to use objective evidence than subjective judgements, for example, if we talk that TMMLU-plus contains questions that require additional images/tables, it is advised to report a percentage maybe within only 50 examples, otherwise this point is not very meaningful. I increase my score to 6 given that I ignore the existence of TMMLU-plus.

Authorsrebuttal2024-06-07

Thanks for the constructive feedback! We will strengthen our objective evidence by providing more relevant samples in the final version.

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC