This article describes the constitution process of the first\nmorpho-syntactically annotated Tunisian Arabish Corpus (TArC). Arabish, also\nknown as Arabizi, is a spontaneous coding of Arabic dialects in Latin\ncharacters and arithmographs (numbers used as letters). This code-system was\ndeveloped by Arabic-speaking users of social media in order to facilitate the\nwriting in the Computer-Mediated Communication (CMC) and text messaging\ninformal frameworks. There is variety in the realization of Arabish amongst\ndialects, and each Arabish code-system is under-resourced, in the same way as\nmost of the Arabic dialects. In the last few years, the focus on Arabic\ndialects in the NLP field has considerably increased. Taking this into\nconsideration, TArC will be a useful support for different types of analyses,\ncomputational and linguistic, as well as for NLP tools training. In this\narticle we will describe preliminary work on the TArC semi-automatic\nconstruction process and some of the first analyses we developed on TArC. In\naddition, in order to provide a complete overview of the challenges faced\nduring the building process, we will present the main Tunisian dialect\ncharacteristics and their encoding in Tunisian Arabish.\n