Prompting
How to Write AI Video Prompts That Give You More Control
A practical six-part framework for AI video prompts, with copyable examples for cinematic scenes, product videos, explainers, and image-to-video.

Most disappointing AI video prompts are not too short. They are unclear about what should happen on screen. “Make it cinematic” leaves the model to decide the subject, shot, movement, light, pace, and even what cinematic means.
A useful prompt works more like a small piece of creative direction. It names the visible ingredients, gives each one a job, and leaves space where you are happy for the model to improvise.
Before writing, choose the input that should carry the most authority. The text-to-video versus image-to-video guide explains when to let a prompt establish the frame and when an approved image should establish it instead.
The six-part AI video prompt framework
Google’s current Veo guidance proposes a five-part formula: cinematography, subject, action, context, and style or ambiance. Runway similarly separates visual description from subject, environmental, and camera motion. The practical overlap is more useful than any model-specific vocabulary.
Use this structure as a checklist, not a sentence you must fill mechanically:
[Shot and camera] of [subject] [visible action] in [setting]. [Lighting, palette, and visual style]. [Subject, scene, and camera timing]. [Sound, if needed].
1. Shot and camera
Decide where the viewer is and how the camera behaves. Start with one clear instruction:
- Framing: close-up, medium shot, wide establishing shot, overhead shot
- Angle: eye level, low angle, high angle, over-the-shoulder
- Movement: locked camera, slow push-in, handheld tracking shot, gentle arc
- Focus: shallow depth of field, deep focus, rack focus from the foreground to the subject
Camera language should describe an observable result. “Premium camera” is vague. “Slow push-in, eye-level medium close-up, shallow depth of field” gives the model something concrete to stage.
2. Subject
Identify the main person, object, or focal point. Include only details that need to stay recognizable: clothing, material, age range, product shape, or defining colors.
Long character biographies rarely help a five-second shot. Translate backstory into details the viewer can actually see. Instead of “an ambitious founder who has struggled for years,” try “a tired founder in a rumpled blue shirt, alone at a desk covered with handwritten notes.”
3. Visible action
Use verbs. Describe the subject’s movement, expression, or interaction in the order it should occur.
“A woman enjoys coffee” asks the model to invent the behavior. “She lifts the cup, takes a small sip, lowers it, and looks toward the window” provides a short, filmable sequence.
4. Setting
Name the environment and the details that affect the scene: time of day, weather, surface, background activity, or spatial relationship.
Keep background instructions subordinate to the main action. If the prompt contains six equally important events, the model has no clear visual hierarchy.
5. Look and light
Choose a small set of compatible visual cues: documentary realism, tactile stop motion, restrained product photography, warm afternoon light, cool fluorescent office light, muted earth tones.
Avoid stacking every aesthetic you like. “Minimalist, maximalist, photoreal, watercolor, retro-futurist documentary” contains competing directions. Two or three aligned choices are usually more controllable.
6. Timing, scene motion, and sound
Describe how motion unfolds: slowly, abruptly, continuously, in a seamless loop, or in a specific sequence. Mention environmental movement where it matters—steam drifting, fabric moving in a breeze, dust following a vehicle.
If the model or workflow generates audio, specify only purposeful sound: quiet cafe room tone, porcelain touching a saucer, distant traffic, no dialogue. Audio direction should support the picture rather than introduce a second story.
Weak prompt versus directed prompt
Consider this starting point:
Make a cinematic video of a woman drinking coffee in a cafe.
It identifies a subject, action, and location, but almost every production choice remains unresolved. That can be useful during exploration. It is not a reliable way to reproduce a particular shot.
Now direct the same idea:
Eye-level medium shot of a woman in a cream sweater seated alone at a small round table beside a cafe window. She lifts a white cup, takes a quiet sip, lowers it to the saucer, and smiles toward the window. The camera performs a very slow push-in. Warm late-afternoon backlight, natural skin tones, soft reflections in the glass, shallow depth of field. Steam drifts from the cup; background patrons remain softly out of focus. Subtle cafe room tone and the sound of porcelain touching the saucer. Continuous shot, calm pacing.

The second prompt is not better because it is longer. It is better because each phrase resolves a production choice. If you do not care how the camera moves, omit that instruction and spend the model’s attention elsewhere.
Prompting a complete video is a different job
A shot prompt directs a few seconds. An end-to-end text-to-video workflow may also write a script, choose scenes, add narration, music, captions, and adapt the result to a platform. It needs a brief above the shot level. If the narration is already approved, start with the more constrained script-to-video workflow.
For a complete video, provide:
- Outcome: what should the viewer understand, feel, or do?
- Audience: who is watching and what do they already know?
- Format: TikTok, Reel, YouTube Short, landscape explainer, product page, or presentation.
- Duration: a real target such as 30 seconds, not simply “short.”
- Narrative shape: hook, problem, demonstration, proof, conclusion, call to action.
- Source of truth: a product URL, document, script, approved claims, images, or brand assets.
- Visual direction: generated scenes, product media, stock footage, presenter, character, or a deliberate mix.
- Audio and text: narrator, language, music energy, caption treatment, and words that must appear exactly.
Create a 30-second vertical product video for independent cafe owners who lose time making staff schedules. The goal is to drive free-trial visits. Open on the frustration of three conflicting shift messages, introduce the scheduling dashboard by second 7, demonstrate drag-and-drop shift planning, then show the finished weekly schedule on a phone. Use our attached product screenshots as the source of truth; do not invent interface elements or performance claims. Calm, capable narration, concise on-screen captions, warm documentary-style workplace footage, and restrained electronic music. End with the exact line: “Build next week’s schedule in minutes.”
Notice that this brief separates claims from creative freedom. The tool can choose transitions and supporting imagery, but it should not invent product capabilities.
Copyable AI video prompt examples
Treat these as starting structures. Replace the subject, setting, action, and constraints instead of pasting them unchanged.
Cinematic establishing shot
Wide establishing shot of a small coastal train emerging from a dark tunnel onto a cliffside track above the sea. The camera glides parallel to the train as morning fog separates into thin ribbons around it. Pale sunrise, cool blue shadows, restrained natural color, fine film grain. Waves break far below; seabirds cross the frame once. Continuous shot with measured, realistic motion and distant rail sound.
Clean product reveal
Macro close-up of a matte black insulated bottle standing on wet slate. A narrow band of light travels slowly from base to lid, revealing condensation and the engraved mark. Locked camera with a subtle focus pull from the water droplets to the logo. Dark neutral studio, crisp edge light, high contrast, precise commercial product photography. One continuous reveal; quiet water drops and a low tonal swell.
Visual explainer
Top-down shot of a clean desk as three paper cards labeled “Research,” “Draft,” and “Review” slide into a simple left-to-right workflow. A hand moves the Draft card forward, and the other cards make room. Soft daylight, off-white paper, charcoal lettering, restrained editorial style. Locked camera, smooth stop-motion movement, no extra labels.
Creator-style opening
Handheld medium close-up of a creator at a kitchen counter holding a colorful running shoe close to the lens. She lowers the shoe, makes direct eye contact, and points to the sole while speaking. Bright natural window light, believable phone-camera exposure, energetic but not frantic movement, everyday home background. Vertical framing with clear space for captions in the lower third.
For image-to-video, prompt the change—not the image
An input image already communicates the subject, composition, lighting, colors, and starting pose. Repeating all of that can compete with the source. Use the text prompt to describe what moves, how the environment reacts, and what the camera does.
Runway’s image-to-video guidance makes this distinction explicit: the image supplies visual information, while the text should focus primarily on motion and temporal progression.
The camera [movement] as the subject [action]. [Environmental response]. [Timing or end state].
For the cafe image above, an image-to-video instruction could be much smaller:
The camera slowly pushes toward the subject as she takes a small sip and lowers the cup. Steam drifts upward and catches the window light. Her expression relaxes into a subtle smile. Background movement remains soft and natural. Continuous, unhurried shot.
Check the source image for motion cues before generating. A parked car surrounded by flying dust already implies speed; asking it to remain perfectly still creates a contradiction the text prompt may not overcome.
Use generation as a controlled conversation
Trying to perfect a prompt before the first result is usually slower than structured iteration. Use this loop:
- Write the simplest complete version of the shot.
- Generate and identify the largest mismatch—not every small flaw.
- Change one category: camera, action, setting, look, or timing.
- Keep the parts that already work worded consistently.
- Save successful prompt-and-result pairs as references for later scenes.
If the subject moves correctly but the shot feels too energetic, change only camera and timing: replace “handheld tracking” with “locked camera” or “slow push-in.” If the framing is right but the action is muddled, shorten the action sequence.
Positive instructions are often clearer than prohibitions. Runway’s Gen-4 guidance, for example, recommends “locked camera; the camera remains still” rather than repeatedly saying “no camera movement.” Describe the desired state whenever you can.
What a stronger prompt cannot fix
Prompting improves direction; it does not guarantee precision. A prompt cannot reliably compensate for:
- a blurry, distorted, or contradictory reference image;
- too many simultaneous actions in a very short clip;
- continuity demands that exceed the model or workflow’s reference controls;
- exact product interfaces when no approved screenshots or footage are supplied;
- unsupported aspect ratios, durations, audio, or editing controls;
- rights you do not have to a person’s likeness, music, brand assets, or source footage.
For factual, product, training, or commercial videos, attach a source of truth and review the script before rendering. Creative detail and factual accuracy are different problems.
Frequently asked questions
How long should an AI video prompt be?
Long enough to resolve the choices you care about. A simple image-to-video motion can work in one sentence. A complete product video needs a broader brief. Length is not the goal; unambiguous direction is.
Should I include every part of the framework?
No. Omitted details give the model room to explore. Include a component when its absence would make the result wrong for your purpose.
Do negative prompts work for video?
Support differs by model. Current Runway guidance recommends positive phrasing for its video models. “Locked camera” is generally clearer than “no camera movement.” Follow the controls and documentation for the model you are actually using.
What is the best prompt for consistent results?
There is no universal phrase that creates consistency. Reuse approved reference images, stable subject descriptions, compatible motion, and the same visual direction. Change one variable at a time so you can tell what affected the output.
Should a prompt describe dialogue word for word?
If exact wording matters, provide it as an approved script or exact on-screen text rather than loose creative direction. Also specify the speaker, language, tone, and whether captions must match verbatim.
Sources and further reading
- Google Cloud: Ultimate prompting guide for Veo 3.1
- Runway Academy: Prompting guide
- Runway: Gen-4 video prompting guide
- Runway: Camera terms, prompts, and examples
This guide was written by the Brevity editorial team from the linked primary documentation and hands-on creative-direction patterns. It describes general prompting practice; individual models and product controls change over time.