0% found this document useful (0 votes)
55 views1 page

Sora: AI Video Generation Overview

Uploaded by

Paolo Molajoni
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
55 views1 page

Sora: AI Video Generation Overview

Uploaded by

Paolo Molajoni
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Menu

Creating video
from text
Sora is an AI model that can create
realistic and imaginative scenes
from text instructions.

Read technical report

All videos on this page were generated


directly by Sora without modification.

Capabilities Safety Research

We’re teaching AI to understand and


simulate the physical world in motion,
with the goal of training models that help
people solve problems that require real-
world interaction.

Introducing Sora, our text-to-video


model. Sora can generate videos up to a
minute long while maintaining visual
quality and adherence to the user’s
prompt.

1 of 9

00:59

Prompt: A stylish woman walks down a Tokyo 0:00 / 0:00 Pro


street filled with warm glowing neon and… ap
animated
more city signage. She wears a black the
mo
leather jacket, a long red dress, and black as
boots, and carries a black purse. She wears dra

Today, Sora is becoming available to red


teamers to assess critical areas for harms
or risks. We are also granting access to a
number of visual artists, designers, and
filmmakers to gain feedback on how to
advance the model to be most helpful for
creative professionals.

We’re sharing our research progress early


to start working with and getting
feedback from people outside of OpenAI
and to give the public a sense of what AI
capabilities are on the horizon.

1 of 8

00:25

Prompt: Historical footage of California 0:00 / 0:00 Pro


during the gold rush. tha
sm
mo
zen
san

Sora is able to generate complex scenes


with multiple characters, specific types of
motion, and accurate details of the
subject and background. The model
understands not only what the user has
asked for in the prompt, but also how
those things exist in the physical world.

1 of 8

00:20

Prompt: The camera follows behind a white 0:00 / 0:00 Pro


vintage SUV with a black roof rack as it… tra
speeds
more up a steep dirt road surrounded by
pine trees on a steep mountain slope, dust
kicks up from it’s tires, the sunlight shines on

The model has a deep understanding of


language, enabling it to accurately
interpret prompts and generate
compelling characters that express
vibrant emotions. Sora can also create
multiple shots within a single generated
video that accurately persist characters
and visual style.

1 of 8

00:20

Prompt: Tour of an art gallery with many 0:00 / 0:00 Pro


beautiful works of art in different styles. bu
bu
mo
en
sh

The current model has weaknesses. It


may struggle with accurately simulating
the physics of a complex scene, and may
not understand specific instances of
cause and effect. For example, a person
might take a bite out of a cookie, but
afterward, the cookie may not have a bite
mark.

The model may also confuse spatial


details of a prompt, for example, mixing
up left and right, and may struggle with
precise descriptions of events that take
place over time, like following a specific
camera trajectory.

1 of 5

00:19

Prompt: Step-printing scene of a person 0:00 / 0:00 Pro


running, cinematic film shot in 35mm. fro
oth
mo
Weakness: Sora sometimes creates gra
physically implausible motion. gra
We
lea
pe
an
ap
pla
co

Safety

We’ll be taking several important safety


steps ahead of making Sora available in
OpenAI’s products. We are working with
red teamers — domain experts in areas
like misinformation, hateful content, and
bias — who will be adversarially testing
the model.

We’re also building tools to help detect


misleading content such as a detection
classifier that can tell when a video was
generated by Sora. We plan to include
C2PA metadata in the future if we deploy
the model in an OpenAI product.

In addition to us developing new


techniques to prepare for deployment,
we’re leveraging the existing safety
methods that we built for our products
that use DALL·E 3, which are applicable to
Sora as well.

For example, once in an OpenAI product,


our text classifier will check and reject
text input prompts that are in violation of
our usage policies, like those that request
extreme violence, sexual content, hateful
imagery, celebrity likeness, or the IP of
others. We’ve also developed robust
image classifiers that are used to review
the frames of every video generated to
help ensure that it adheres to our usage
policies, before it’s shown to the user.

We’ll be engaging policymakers,


educators and artists around the world to
understand their concerns and to identify
positive use cases for this new
technology. Despite extensive research
and testing, we cannot predict all of the
beneficial ways people will use our
technology, nor all the ways people will
abuse it. That’s why we believe that
learning from real-world use is a critical
component of creating and releasing
increasingly safe AI systems over time.

1 of 10

00:09

Prompt: The camera directly faces colorful 0:00 / 0:00 Pro


buildings in Burano Italy. An adorable… sta
dalmation
more looks through a window on a life
mo
building on the ground floor. Many people wa
are walking and cycling along the canal ren

Research techniques

Sora is a diffusion model, which


generates a video by starting off with one
that looks like static noise and gradually
transforms it by removing the noise over
many steps.

Sora is capable of generating entire


videos all at once or extending generated
videos to make them longer. By giving
the model foresight of many frames at a
time, we’ve solved a challenging problem
of making sure a subject stays the same
even when it goes out of view temporarily.

Similar to GPT models, Sora uses a


transformer architecture, unlocking
superior scaling performance.

We represent videos and images as


collections of smaller units of data called
patches, each of which is akin to a token
in GPT. By unifying how we represent
data, we can train diffusion transformers
on a wider range of visual data than was
possible before, spanning different
durations, resolutions and aspect ratios.

Sora builds on past research in DALL·E


and GPT models. It uses the recaptioning
technique from DALL·E 3, which involves
generating highly descriptive captions
for the visual training data. As a result,
the model is able to follow the user’s text
instructions in the generated video more
faithfully.

In addition to being able to generate a


video solely from text instructions, the
model is able to take an existing still
image and generate a video from it,
animating the image’s contents with
accuracy and attention to small detail.
The model can also take an existing video
and extend it or fill in missing frames.
Learn more in our technical report.

Sora serves as a foundation for models


that can understand and simulate the real
world, a capability we believe will be an
important milestone for achieving AGI.

Research Leads
Bill Peebles & Tim Brooks

Systems Lead
Connor Holmes

Contributors
Clarence Wing Yin Ng
David Schnurr
Eric Luhman
Joe Taylor
Li Jing
Natalie Summers
Ricky Wang
Rohan Sahai
Ryan O’Rourke
Troy Luhman
Will DePue
Yufei Guo

Special Thanks
Bob McGrew, Brad Lightcap, Chad Nelson,
David Medina, Gabriel Goh, Greg Brockman, Ian Sohl,
Jamie Kiros, James Betker, Jason Kwon,
Hannah Wong, Mark Chen, Michelle Fradin,
Mira Murati, Nick Turley, Prafulla Dhariwal,
Rowan Zellers, Sarah Yoo, Sandhini Agarwal,
Sam Altman, Srinivas Narayanan & Wesam Manassra

Communications
Elie Georges
Justin Wang
Kendra Rimbach
Niko Felix
Thomas Degry
Veit Moeller

Legal
Che Chang
Fred von Lohmann
Gideon Myles
Tom Stasi

External Engagement
Alex Baker-Whitcomb, Allie Teague, Anna Makanju,
Anna McKean, Becky Waite, Brittany Smith,
Chan Park, Chris Lehane, David Duxin,
David Robinson, James Hairston, Jonathan Lachman,
Justin Oswald, Krithika Muthukumar, Lane Dilg,
Leher Pathak, Ola Nowicka, Ryan Biddy,
Sandro Gianella, Stephen Petersilge, Tom Rubin &
Varun Shetty

Executive Producer
Aditya Ramesh

Built by OpenAI in San Francisco, California


Published February 15, MMXXIV

Research API
Overview Overview
Index Pricing
GPT-4 Docs
DALL·E 3
Sora

ChatGPT Company
Overview About
Team Blog
Enterprise Careers
Pricing Charter
Try ChatGPT Security
Customer stories
Safety

OpenAI © 2015 – 2024 Social


Terms & policies Twitter
Privacy policy YouTube
Brand guidelines GitHub
SoundCloud
LinkedIn

Back to top

Common questions

Powered by AI

Real-world interaction is significant in AI models like Sora because it represents a crucial step toward developing systems that can operate and adapt in dynamic environments, reflecting OpenAI's vision of achieving AGI. By understanding and simulating the physical world, models like Sora can solve complex real-world problems, enhancing their applicability across diverse industries and contributing to the evolution of interactive AI technologies .

The main objective of introducing the Sora text-to-video model by OpenAI is to train models that help people solve problems requiring real-world interaction. Sora aims to generate videos accurately based on text prompts, maintaining visual quality and adherence to user instructions. The introduction of Sora serves to understand and simulate the physical world in motion, contributing greatly to OpenAI's goal of advancing artificial general intelligence (AGI).

OpenAI addresses ethical concerns associated with Sora's deployment by employing mechanisms like adversarial testing through red teaming, detection classifiers to identify generated content, and text input filters to prevent inappropriate requests. The company also engages with policymakers, educators, and artists globally to understand concerns and develop positive use cases, ensuring the technology is deployed ethically and responsibly .

Sora builds on past research in DALL·E and GPT models, using the transformer architecture to facilitate superior scaling performance. It employs recaptioning techniques from DALL·E 3, creating highly descriptive captions for visual training data, ensuring more faithful adherence to text instructions in generated videos. Additionally, Sora uses diffusion model techniques to generate videos from static noise, progressively refining the output to achieve reliable video generation .

Recaptioning techniques from DALL·E 3 are incorporated into Sora's video generation process by generating highly descriptive captions for visual training data, enhancing the model's understanding of text instructions. This enables Sora to more accurately interpret user prompts and translate them into faithful visual representations, improving the correlation between instruction and video output .

Sora faces challenges in accurately simulating the physics of complex scenes and understanding cause and effect, such as inconsistencies in action outcomes like missing bite marks on objects. It also struggles with accurately interpreting spatial details, potentially confusing left and right, and maintaining precise descriptions over time. These challenges impact its usability by affecting the realistic portrayal of actions and scene continuity, which can limit its application in scenarios requiring detailed and precise animation .

Sora's use of transformer architecture allows for superior scaling performance, influencing its ability to handle a wide range of visual data differing in duration, resolutions, and aspect ratios. This is a significant improvement over previous models, as it enables the processing of visual information similarly to text data in GPT models, allowing for a more coherent and consistent output across varied video scenarios .

OpenAI plans to ensure the safe use of the Sora model by working with red teamers who are domain experts tasked with adversarial testing for misinformation, hateful content, and bias. The company is developing detection classifiers to identify videos generated by Sora and employing text classifiers to reject prompts that violate usage policies. OpenAI also reviews each video frame with robust classifiers to ensure compliance with safety policies before they are shown to users .

Sora accommodates the creative needs of professionals by offering features like complex scene generation, accurate motion depiction, and the ability to maintain visual style consistency across multiple shots. It enables these professionals to depict intricate narratives and dynamic visuals with ease, increasing creative possibilities without traditional production constraints .

The potential benefits of using Sora include its ability to create realistic and imaginative scenes from text instructions, enabling artists and creators to visually express ideas without conventional production means. However, its limitations include challenges in simulating physics accurately and potential confusion with spatial or temporal details. These limitations may restrict its effectiveness in scenarios where precise and consistent depiction of events is crucial .

You might also like