Emu Video

Video generation

Video generation from text prompts

Emu Video screenshot

About Emu Video

Emu Video is a cutting-edge tool designed for text-to-video generation through explicit image conditioning. By utilizing diffusion models, Emu Video breaks down the generation process into two distinct steps: first, creating an image based on a text prompt, and then generating a video using both the prompt and the generated image.This unique approach allows for the efficient training of top-tier video generation models. Setting itself apart from previous methods that rely on a deep cascade of models, Emu Video only requires two diffusion models to produce high-quality 512px, 4-second-long videos at 16fps. When compared to other models like Make-a-Video (MAV), Imagen-Video (IMAGEN), and Align Your Latents (AYL), Emu Video consistently delivers state-of-the-art results in text-to-video generation.Human raters have recognized Emu Video's 512-pixel, 16fps, 4-second-long videos as the most compelling in terms of quality and fidelity to the provided prompt. The tool was developed by a team of experts including Rohit Girdhar, Mannat Singh, Andrew Brown, and others, with equal technical contributions from Girdhar and Singh.Emu Video extends its gratitude to various collaborators who contributed to the project by providing data and infrastructure support. Additionally, the tool upholds strict privacy and cookie policies, which can be reviewed on its official website.

Companies Are Making AI Skills Mandatory

Performance reviews and hiring now depend on AI proficiency

Meta
Shopify
Microsoft
Duolingo
Klarna
Google
Opendoor
Thomson Reuters
Fiverr
Amazon

Track the Impact of Your AI Usage

Document your productivity gains and build your AI portfolio for performance reviews

Start Tracking Free