This paper presents a multimodal framework integrating generative artificial intelligence (AI) into film production and educational content creation, with a focus on cultural adaptability. Combining text generation, visual synthesis, audio-video integration, and interactive Q&A modules, the framework enables automated production of culturally enriched educational media, using Guangdong culture as a case study. A demonstration using ChatGPT-4o, Stable Diffusion XL, and Pika Labs explored a futuristic urban scenario, with evaluations from 15 participants highlighting strong engagement, narrative depth, and cultural resonance. Comparative tool analysis revealed trade-offs in generation speed, quality, and post-editing flexibility, while user feedback emphasized the importance of fine-grained control, educational clarity, and ethical safeguards. This study offers practical and theoretical insights for advancing culturally responsive, AI-driven educational media at the intersection of artificial intelligence, creative industries, and education.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex