Narrative cloze tasks are widely used to study whether a model understands a story well enough to pick a coherent continuation. A recurring problem is that such tasks can be partly solved without reading the story at all: incorrect endings often differ from correct ones in style or wording, so a model can succeed by reading the ending alone. This work treats that shortcut not as a nuisance to be hidden, but as the quantity to measure. We build an automatic narrative cloze task from public ROCStories-style data in which both endings are genuinely human-written: the wrong ending is retrieved from a different human story by nearest-neighbor similarity, rather than generated from a template. On this task we report three transparent baselines — an endings-only classifier, a context-plus-ending classifier, and a compressed bag-of-words representation — and measure each one's gain over the endings-only shortcut. Across four random seeds (20,000 examples each), the endings-only classifier already reaches 58.9%, confirming the shortcut is real. The context-plus-ending model adds +5.2 points and the compressed representation adds +8.0 points over it. Under an entity-masking stress test that removes name cues, both gains shrink to roughly +3.9 points, showing that a meaningful part of the apparent "context understanding" is entity-string overlap. We make no claim of causal or world-model reasoning. The contribution is a reproducible, shortcut-controlled evaluation protocol and an honest accounting of how much measured context gain survives when surface cues are removed.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex