Automated Collection of Evaluation Dataset for Semantic Search in Low-Resource Domain Language

Domain-specific languages that use a lot of specific terminology often fall\ninto the category of low-resource languages. Collecting test datasets in a\nnarrow domain is time-consuming and requires skilled human resources with\ndomain knowledge and training for the annotation task. This study addresses the\nchallenge of automated collecting test datasets to evaluate semantic search in\nlow-resource domain-specific German language of the process industry. Our\napproach proposes an end-to-end annotation pipeline for automated query\ngeneration to the score reassessment of query-document pairs. To overcome the\nlack of text encoders trained in the German chemistry domain, we explore a\nprinciple of an ensemble of "weak" text encoders trained on common knowledge\ndatasets. We combine individual relevance scores from diverse models to\nretrieve document candidates and relevance scores generated by an LLM, aiming\nto achieve consensus on query-document alignment. Evaluation results\ndemonstrate that the ensemble method significantly improves alignment with\nhuman-assigned relevance scores, outperforming individual models in both\ninter-coder agreement and accuracy metrics. These findings suggest that\nensemble learning can effectively adapt semantic search systems for\nspecialized, low-resource languages, offering a practical solution to resource\nlimitations in domain-specific contexts.\n

Paper

References (25)

Scroll for more · 13 remaining

Similar papers

© 2026 NYSGPT2525 LLC