Visual Narratives: Bridging Images and Stories Using NLP

In today’s visual world, the power of turning images into narrative experiences is still an untapped market. The problem is that there lacks a simple yet powerful system to automatically translate images into engaging, contextually rich stories — until now this has been the goal of "Image to Story Narration" work. The system uses state-of-the-art AI algorithms based on the deep learning and modern natural-language processing mechanisms to comprehend both textual input data from users as well visual content in order to build consistent, creative stories. It has an easy to use interface where users can upload images and pick a genre for their stories such as Horror, Action, Romance, Comedy, Historical or Sci-fi.The image-to-caption module uses deep learning to generate textual descriptions and the caption-to-story-module employs a GPT-2 model. This twin-module design guarantees tales that are contextually perfect yet also imaginative compelling. Preliminary results suggest the system is competitive in generating a variety of engaging narratives, spanning multiple genres. The resulting stories, both in terms of accuracy and creativity outperform plain factual captions showing a high level satisfaction towards the successful translation from visual to narrative form. In early testing, the system has been adept at producing original science fiction stories and highly factual historic tales. Applications for this "Image to Story Narration" work are broad, from improving sighted multimedia teaching aids and presentations to creating immersive storytelling experiences for the blind. The proposed system is very effective at producing relevant sentences for images. It also generates descriptions that are notably more true to the specific image content than previous work.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC