<scp>MotionBlend GAN</scp> : An Approach for Realistic Video Content Creation With Embedded Approach for Human Subjects
ABSTRACT The proposed MotionBlend GAN model marks a significant step forward in video synthesis by blending the motion from a source video with the appearance of a target person's image. As training progresses, the model improves video creation by enhancing the smoothness and natural flow of motion, resulting in more coherent and lifelike videos. Using advanced techniques like MoBConv blocks of EfficientNet‐B7, OpenPose for precise pose detection, ResNet blocks for feature integration, and a 3D CNN discriminator, the model produces high‐quality videos that maintain both spatial and temporal consistency. After 200 epochs, the model achieved an adversarial loss of 0.2265, with metrics like PSNR at 20.246, SSIM at 0.867, and LPIPS at 0.178. The high PSNR and SSIM values, along with the low LPIPS, show that the generated frames are well aligned and preserve important details. These results highlight the model's strong performance over time, consistently generating visually convincing videos of human activities using a reference image and source video. The model effectively transfers motion from video to image, creating realistic videos of human activity in comparison to existing models.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex