HRE-LSC: A Hyper-Relational Data Enhancement Framework for Long Tail Distribution and Structural Consistency
Hyper-relational extraction aims to identify complex factual structures from unstructured text, where each instance consists of a core triplet and multiple qualified attributes. However, existing hyper-relational datasets often suffer from long-tail relation distributions and pseudo-negative samples, which limit the performance of hyper-relational extraction models. Although large language models (LLMs) provide new opportunities for data augmentation, they may introduce semantically inconsistent relations and structurally unreliable samples in complex scenarios, reducing the quality of generated data. To address these challenges, this paper proposes a structure-consistency-driven semi-automatic data augmentation framework, termed HRE-LSC. The framework consists of two key components: (1) a distribution-aware generation strategy that selectively generates samples for low-frequency relations according to relation frequency distributions, thereby alleviating data imbalance; and (2) a hierarchical logical consistency verification mechanism based on natural language inference (NLI), which progressively verifies the consistency of core triplets and qualified attributes to filter unreliable generated samples and reduce pseudo-negative samples. Experiments on the HyperRED dataset demonstrate that the proposed method improves the F1 score of Text2NKG from 83.81% to 84.7%, with more significant improvements observed on long-tail relation subsets. The results indicate that the proposed framework effectively enhances the quality and structural consistency of generated hyper-relational data while mitigating the effects of long-tail distributions and pseudo-negative samples without requiring additional manual annotations.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex