Emu Video
Video generation from text prompts
About Emu Video
Emu Video is a cutting-edge tool designed for text-to-video generation through explicit image conditioning. By utilizing diffusion models, Emu Video breaks down the generation process into two distinct steps: first, creating an image based on a text prompt, and then generating a video using both the prompt and the generated image.This unique approach allows for the efficient training of top-tier video generation models. Setting itself apart from previous methods that rely on a deep cascade of models, Emu Video only requires two diffusion models to produce high-quality 512px, 4-second-long videos at 16fps. When compared to other models like Make-a-Video (MAV), Imagen-Video (IMAGEN), and Align Your Latents (AYL), Emu Video consistently delivers state-of-the-art results in text-to-video generation.Human raters have recognized Emu Video's 512-pixel, 16fps, 4-second-long videos as the most compelling in terms of quality and fidelity to the provided prompt. The tool was developed by a team of experts including Rohit Girdhar, Mannat Singh, Andrew Brown, and others, with equal technical contributions from Girdhar and Singh.Emu Video extends its gratitude to various collaborators who contributed to the project by providing data and infrastructure support. Additionally, the tool upholds strict privacy and cookie policies, which can be reviewed on its official website.