Exploiting Class Labels to Boost Performance on Embedding-based Text Classification

Text classification is one of the most frequent tasks for processing textual\ndata, facilitating among others research from large-scale datasets. Embeddings\nof different kinds have recently become the de facto standard as features used\nfor text classification. These embeddings have the capacity to capture meanings\nof words inferred from occurrences in large external collections. While they\nare built out of external collections, they are unaware of the distributional\ncharacteristics of words in the classification dataset at hand, including most\nimportantly the distribution of words across classes in training data. To make\nthe most of these embeddings as features and to boost the performance of\nclassifiers using them, we introduce a weighting scheme, Term\nFrequency-Category Ratio (TF-CR), which can weight high-frequency,\ncategory-exclusive words higher when computing word embeddings. Our experiments\non eight datasets show the effectiveness of TF-CR, leading to improved\nperformance scores over the well-known weighting schemes TF-IDF and KLD as well\nas over the absence of a weighting scheme in most cases.\n

Paper

References (33)

Scroll for more · 21 remaining

Similar papers

© 2026 NYSGPT2525 LLC