Systems for story generation are asked to produce plausible and enjoyable\nstories given an input context. This task is underspecified, as a vast number\nof diverse stories can originate from a single input. The large output space\nmakes it difficult to build and evaluate story generation models, as (1)\nexisting datasets lack rich enough contexts to meaningfully guide models, and\n(2) existing evaluations (both crowdsourced and automatic) are unreliable for\nassessing long-form creative text. To address these issues, we introduce a\ndataset and evaluation platform built from STORIUM, an online collaborative\nstorytelling community. Our author-generated dataset contains 6K lengthy\nstories (125M tokens) with fine-grained natural language annotations (e.g.,\ncharacter goals and attributes) interspersed throughout each narrative, forming\na robust source for guiding models. We evaluate language models fine-tuned on\nour dataset by integrating them onto STORIUM, where real authors can query a\nmodel for suggested story continuations and then edit them. Automatic metrics\ncomputed over these edits correlate well with both user ratings of generated\nstories and qualitative feedback from semi-structured user interviews. We\nrelease both the STORIUM dataset and evaluation platform to spur more\nprincipled research into story generation.\n
Paper
References (36)
Scroll for more · 24 remaining