Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture

This study introduces a comprehensive benchmark designed to evaluate the performance of large language models (LLMs) in understanding and processing cultural knowledge, with a specific focus on Hakka culture as a case study. Leveraging Bloom's Taxonomy, the study develops a multi-dimensional framework that systematically assesses LLMs across six cog-nitive domains: Remembering, Understanding, Applying, Analyzing, Evaluating, and Creating. This benchmark ex-tends beyond traditional single-dimensional evaluations by providing a deeper analysis of LLMs' abilities to handle cul-turally specific content, ranging from basic recall of facts to higher-order cognitive tasks such as creative synthesis. Ad-ditionally, the study integrates Retrieval-Augmented Generation (RAG) technology to address the challenges of minority cultural knowledge representation in LLMs, demonstrating how RAG enhances the models' performance by dynamically incorporating relevant external information. The results high-light the effectiveness of RAG in improving accuracy across all cognitive domains, particularly in tasks requiring precise retrieval and application of cultural know ledge. However, the findings also reveal the limitations of RAG in creative tasks, underscoring the need for further optimization. This benchmark provides a robust tool for evaluating and com-paring LLMs in culturally diverse contexts, offering valuable insights for future research and development in AI -driven cultural knowledge preservation and dissemination.

Paper

References (23)

Scroll for more · 11 remaining

Similar papers

© 2026 NYSGPT2525 LLC