Language generation tasks that seek to mimic human ability to use language\ncreatively are difficult to evaluate, since one must consider creativity,\nstyle, and other non-trivial aspects of the generated text. The goal of this\npaper is to develop evaluation methods for one such task, ghostwriting of rap\nlyrics, and to provide an explicit, quantifiable foundation for the goals and\nfuture directions of this task. Ghostwriting must produce text that is similar\nin style to the emulated artist, yet distinct in content. We develop a novel\nevaluation methodology that addresses several complementary aspects of this\ntask, and illustrate how such evaluation can be used to meaningfully analyze\nsystem performance. We provide a corpus of lyrics for 13 rap artists, annotated\nfor stylistic similarity, which allows us to assess the feasibility of manual\nevaluation for generated verse.\n