CAp 2017 challenge: Twitter Named Entity Recognition

The paper describes the CAp 2017 challenge. The challenge concerns the problem of Named Entity Recognition (NER) for tweets written in French. We first present the data preparation steps we followed for constructing the dataset released in the framework of the challenge. We begin by demonstrating why NER for tweets is a challenging problem especially when the number of entities increases. We detail the annotation process and the necessary decisions we made. We provide statistics on the inter-annotator agreement, and we conclude the data description part with examples and statistics for the data. We, then, describe the participation in the challenge, where 8 teams participated, with a focus on the methods employed by the challenge participants and the scores achieved in terms of F$_1$ measure. Importantly, the constructed dataset comprising $\sim$6,000 tweets annotated for 13 types of entities, which to the best of our knowledge is the first such dataset in French, is publicly available at \url{this http URL} .

Paper

References (9)

05Detection des entitees nommees dans le cadre de la conference cap2017 · CAP
06Alexsandro Fonseca, Fatma Mallek, Billal Belainine, Fatiha Sadat, and Dien Dinh. Reconnaissance des entités nommées dans les messages twitter en français2017
07Amu-lif at cap 2017 ner challenge: how generic are ne taggers to new tagsets and new language registers ? In CAP2017
08Reconnais-sance des entités nommées dans les messages twitter en français2017 · CAP
09Fner-bgru-crf at cap 2017 ner challenge: Bidirectional gru-crf for french named entity recognition in tweets2017 · CAP

Similar papers

© 2026 NYSGPT2525 LLC