Batteries, camera, action! Learning a semantic control space for expressive robot cinematography
Aerial vehicles are revolutionizing the way film-makers can capture shots of\nactors by composing novel aerial and dynamic viewpoints. However, despite great\nadvancements in autonomous flight technology, generating expressive camera\nbehaviors is still a challenge and requires non-technical users to edit a large\nnumber of unintuitive control parameters. In this work, we develop a\ndata-driven framework that enables editing of these complex camera positioning\nparameters in a semantic space (e.g. calm, enjoyable, establishing). First, we\ngenerate a database of video clips with a diverse range of shots in a\nphoto-realistic simulator, and use hundreds of participants in a crowd-sourcing\nframework to obtain scores for a set of semantic descriptors for each clip.\nNext, we analyze correlations between descriptors and build a semantic control\nspace based on cinematography guidelines and human perception studies. Finally,\nwe learn a generative model that can map a set of desired semantic video\ndescriptors into low-level camera trajectory parameters. We evaluate our system\nby demonstrating that our model successfully generates shots that are rated by\nparticipants as having the expected degrees of expression for each descriptor.\nWe also show that our models generalize to different scenes in both simulation\nand real-world experiments. Data and video found at:\nhttps://sites.google.com/view/robotcam.\n
Paper
References (51)
Scroll for more · 38 remaining