Character-Centric Storytelling

Sequential vision-to-language or visual storytelling has recently been one of the areas of focus in computer vision and language modeling domains. Though existing models generate narratives that read subjectively well, there could be cases when these models miss out on generating stories that account and address all prospective human and animal characters in the image sequences. Considering this scenario, we propose a model that implicitly learns relationships between provided characters and thereby generates stories with respective characters in scope. We use the VIST dataset for this purpose and report numerous statistics on the dataset. Eventually, we describe the model, explain the experiment and discuss our current status and future work.

Paper

References (12)

09Computer vision and natural language processing: Recent approaches in multimedia and robotics2016 · ACM Comput. Surv., 49(4):71:1–71:44.
10Figure 7: VIST test split image sequences with stories generated by the character-centric storytelling model
11there are many types of horns and bone . this is a ram horn and it has been shined well . next there is a horn that has been made into a pipe . a curved horn has been rubbed and shined
12Character-centric model): the crowd gathered for the awards ceremony . the speaker gave a great speech . the director gave a brief speech

Similar papers

© 2026 NYSGPT2525 LLC