Deep Fake Creation and Detection of Multimedia Data Using Deep Learning
The task of creating personalized photo realistic talking head models, i.e., systems that synthesize video-sequences of speech expressions mimics of a particular individual is considered.In this work, we present a system for creating talking head models from a handful of photographs (so-called few shot learning) and with limited training time. In fact, our system can generate a reasonable result based on a single photograph (one-shot learning), while adding a few more photographs increases the field it personalization. Similarly, the talking heads created by our model are deep ConvNets that synthesize video frames in a direct manner by a sequence of convolutional operations rather than by warping.We present a system with such few- shot capability. It performs lengthy meta-learning on a large dataset of videos, and after that is able to frame few- and one-shot learning of neural talking head models of previously unseen people as adversarial training problems with high capacity generators and discriminators. The system is able to initialize the parameters of both the generator and the discriminator in a person- specific way, so that training can be based on just a few images and done quickly, despite the need to tune tens of millions of parameters. We show that such an approach is able to learn highly realistic and personalized talking head models of new people and even portrait paintings.