We release edgar-corpus, a novel corpus comprising annual reports from all the publicly traded companies in the us spanning a period of more than 25 years. To the best of our knowledge, edgar-corpus is the largest financial nlp corpus available to date. All the reports are downloaded, split into their corresponding items (sections), and provided in a clean, easy-to-use json format. We use edgar-corpus to train and release edgar-w2v, which are word2vec embeddings for the financial domain. We employ these embeddings in a battery of financial nlp tasks and showcase their superiority over generic glove embeddings and other existing financial word embeddings. We also open-source edgarcrawler, a toolkit that facilitates downloading and extracting future annual reports.