Etutor: An End-to-End Multimodal Framework for AI-Driven Educational Video Generation

Online education platforms often lack personalized learning experiences and effective emotional interaction, partly because existing technologies struggle to generate audio and video synchronously, leading to inconsistencies in semantics and emotional expression. To address this challenge, this paper introduces Etutor, an innovative, end-to-end multimodal AI framework for educational video generation that can rapidly create high-quality, personalized instructional videos from a single text input. At the core of Etutor is a three-layer architecture—encompassing semantics, expression, and technology—coordinated by a Large Language Model (LLM). This architecture uniformly interprets user intent, generates instructional content, and orchestrates all modal parameters to ensure a high degree of semantic and emotional consistency. The framework integrates two key technical modules: NMCTTS, a speech synthesis system optimized for Science, Technology, Engineering, and Mathematics (STEM) content that accurately handles professional terminology and formulas; and CF-RTVideo, a video generation system that significantly enhances the clarity of facial details, the vividness of expressions, and the preservation of identity. To comprehensively evaluate system performance, we introduced the Teaching Cognitive Load Synergy Index (TCLSI) as a new evaluation metric. Experimental results show that Etutor achieved a high score of 0.78 on the TCLSI, outperforming all baseline systems that use separate modular components. This result demonstrates Etutor's comprehensive advantage in balancing instructional content delivery, multimodal expressiveness, and learner cognitive load, offering an effective automated solution to the challenges in modern online education.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC