The experimental landscape in natural language processing for social media is\ntoo fragmented. Each year, new shared tasks and datasets are proposed, ranging\nfrom classics like sentiment analysis to irony detection or emoji prediction.\nTherefore, it is unclear what the current state of the art is, as there is no\nstandardized evaluation protocol, neither a strong set of baselines trained on\nsuch domain-specific data. In this paper, we propose a new evaluation framework\n(TweetEval) consisting of seven heterogeneous Twitter-specific classification\ntasks. We also provide a strong set of baselines as starting point, and compare\ndifferent language modeling pre-training strategies. Our initial experiments\nshow the effectiveness of starting off with existing pre-trained generic\nlanguage models, and continue training them on Twitter corpora.\n