1. Text to Video
July 29, 2026
In one sentence.You start with a written idea: the prompt must build the entire scene.
Text to Video: build the shot from scratch
1
Your idea
The subject, setting, and intended action
2
Your prompt
Shot, action, camera, lighting, and mood
3
The video
A fully generated scene
The text provides all the visual information needed to create the shot.
When to use this mode
- Create a scene without a starting image.
- Design an advertisement, atmosphere, setting, or cinematic shot.
- Explore a visual concept quickly before producing reference assets.
What your prompt should specify
- The framing or point of view.
- The main subject and the useful details of its appearance.
- One observable action.
- The setting and its essential elements.
- One dominant camera move.
- Lighting, mood, and optional audio.
Quick formula
Shot
Subject
Action
Setting
Camera
Lighting / mood
COPY-READY EXAMPLE - ABOUT 8 SECONDS
"Cinematic medium shot of a young architect presenting a scale model in a bright contemporary meeting room. She points to the model while two colleagues listen. The camera tracks slowly from left to right as warm daylight enters through tall windows, creating a calm professional atmosphere. Subtle room ambience."
Why it works: one main subject, one readable action, one camera move, and concrete lighting.
What to avoid and how to improve it
| Avoid | Improved version |
|---|---|
| Create an amazing premium futuristic video with many camera moves and lots of action. | Wide shot of a sleek electric car crossing a quiet desert road at sunrise. The camera follows from a low rear angle while warm light reflects across the bodywork. |
Capabilities and limitations
- Well suited to ideation and entirely new scenes.
- Less suitable when the exact appearance of a person, product, or setting must be preserved.
- Short scenes do not handle long sequences of actions well.
- Perfectly legible text or logos may require a reference image or post-production finishing.
Quick checklist
The main subject is clearly identifiable.
The action can be completed within the selected duration.
The camera follows only one dominant path.
Lighting and mood are described through visible details.
