We present a fairly large, Potential Idiomatic Expression (PIE) dataset for\nNatural Language Processing (NLP) in English. The challenges with NLP systems\nwith regards to tasks such as Machine Translation (MT), word sense\ndisambiguation (WSD) and information retrieval make it imperative to have a\nlabelled idioms dataset with classes such as it is in this work. To the best of\nthe authors' knowledge, this is the first idioms corpus with classes of idioms\nbeyond the literal and the general idioms classification. In particular, the\nfollowing classes are labelled in the dataset: metaphor, simile, euphemism,\nparallelism, personification, oxymoron, paradox, hyperbole, irony and literal.\nWe obtain an overall inter-annotator agreement (IAA) score, between two\nindependent annotators, of 88.89%. Many past efforts have been limited in the\ncorpus size and classes of samples but this dataset contains over 20,100\nsamples with almost 1,200 cases of idioms (with their meanings) from 10 classes\n(or senses). The corpus may also be extended by researchers to meet specific\nneeds. The corpus has part of speech (PoS) tagging from the NLTK library.\nClassification experiments performed on the corpus to obtain a baseline and\ncomparison among three common models, including the BERT model, give good\nresults. We also make publicly available the corpus and the relevant codes for\nworking with it for NLP tasks.\n