Text to Video Prompt Examples You Can Copy and Adapt

Summary

Strong text to video prompt examples follow one skeleton: camera move, scene, a single action, then light and mood. Keep each clip to 5-10 seconds, write what you want to see instead of what to avoid, and change one element per re-run. This guide gives copy-ready prompts for social hooks, product shots, B-roll and talking clips, plus the mistakes to drop.

Dark editing desk at night with a monitor showing a rain-soaked neon street frame and a shot list notebook

These text to video prompt examples are built to be copied, then bent to your own shot. Each one follows the same skeleton: camera move, scene, one action, a few details. Keep it to a single action that fits in 5 to 10 seconds, describe what you want to see rather than what to avoid, and change one element per re-run. The examples below cover hooks, product shots, B-roll, talking-head intros and training clips.

Why do most text to video prompts come back as mush?

Because they read like a mood board. "Epic cinematic masterpiece, 8K, ultra detailed" tells the model nothing it can move. A video model needs physics: who is doing what, where, with which camera behaviour.

Runway's own prompting guide boils it down to a four-part order: camera movement, scene, action, details, with one action per clip. Kling's guide lands on the same ingredients: subject, action, setting, camera language, lighting and mood.

Two habits kill clips. Stacking five actions into one prompt, and writing "then" as if the model had a timeline. It does not. Each generation is one beat.

Storyboard of four blank frames on a desk with a pencil and a toy camera dolly

What is the prompt skeleton that works across models?

Use this order and stop adding words when the shot is clear:

[Shot and camera move]. [Scene and setting]. [Subject doing one action]. [Light, lens, mood].

Here is the same idea three times, from weak to usable.

Weak: "A beautiful cinematic video of a woman in a cafe."

Better: "A young woman lifts a latte to her lips in a sunlit cafe."

Shot-ready: "Slow push-in from medium shot to close-up. Sunlit corner cafe, steam rising from a ceramic cup. A young woman in a cream wool sweater lifts the cup and smiles faintly. Shallow depth of field, soft golden morning light, muted warm tones."

The third version names the camera, the action and the light. Nothing in it asks the model to guess a feeling. Production scale starts here: a prompt you can reuse is a prompt with slots.

Which prompt examples work for social hooks and ad-style clips?

Short, loud, one idea. Vertical framing goes in the prompt, not in your hopes. These are written for 9:16 and 5 to 8 seconds.

Fast whip pan to a cluttered desk at midnight. A laptop screen glows, papers scattered, a cold coffee. A hand slams the laptop shut, then the frame holds still. Handheld, harsh overhead light, 9:16 vertical.

Locked-off macro shot. A single drop of ink falls into clear water and blooms into a spiral. Slow motion, black background, studio lighting, sharp focus on the bloom.

Low-angle tracking shot beside a runner's shoes hitting wet pavement at dawn. Mist at ankle height, street lamps fading out. Smooth stabilised movement, cool blue light.

Two more for creators who publish on a schedule, both for faceless channels:

Slow dolly forward through a dim library aisle. Dust floats in a single shaft of light from a high window. A leather book slides off a shelf and lands open on a reading table. Warm tungsten tones, 16:9.

Smooth crane up from a laptop on a desk to reveal a city skyline at blue hour through a wide window. A mug sits beside the keyboard, the screen reflecting soft light. Calm, steady, documentary feel.

The pattern: a verb the model can render, a texture it can hold (wet pavement, ink, glass), and a camera behaviour that does not fight the subject. Skip the on-screen text. Text inside generated video still drifts and garbles in most models, so add captions in your editor.

How do you prompt product shots and B-roll?

Product work rewards restraint. The model keeps shape and light well when little else is moving.

Slow push-in on a ceramic coffee cup as steam rises, quiet kitchen morning, shallow depth of field, soft window light from the left.

Overhead shot, a pair of hands unboxes a matte black speaker on a pale oak table. The lid lifts, tissue paper settles. Neutral tones, diffused light, no camera shake.

Orbit shot around a perfume bottle on a velvet pedestal. Light glints along the glass edge. Slow, even rotation, dark background, subtle reflection on the surface.

Miniature camera on a slider rail pushing in toward a steaming latte cup on a model cafe set

For B-roll, think in establishing, detail and cutaway. One prompt per type:

Generate three of each and you have a sequence, not a lucky clip. That is also how you build a library: label the winners by shot type and reuse the skeleton.

How do you prompt training and explainer clips?

L&D and explainer work is mostly illustration: a process, a place, a before and after. The prompt job is to show one idea per clip so the voiceover can carry the rest.

Static medium shot of a warehouse worker scanning a box with a handheld scanner, then placing it on a shelf. Bright, even industrial lighting, clean floor, no other movement in frame.

Top-down shot of hands arranging sticky notes into three columns on a whiteboard. Notes are blank, colors yellow, blue and pink. Steady camera, soft daylight, shallow shadows.

Slow pan across a tidy open-plan office at 9 a.m., desks half occupied, monitors showing generic dashboards. Natural light, neutral palette, calm pace.

Notice what is missing: logos, readable screens, signage, names. Anything the viewer has to read should be added in post, where you control the spelling. For a 40-module series, this is the cheapest habit to adopt. Write each prompt as a reusable scene card, tag it with the module it serves, and you will never regenerate a generic "office" shot twice.

What changes when you write prompts for people who talk?

Text to video prompt examples that include dialogue are where results split by model. Some current models generate synchronized audio and lip movement; others give you a silent clip and a face that drifts after a few seconds.

If your model supports speech, be literal and short:

Medium close-up, static camera. A woman at a kitchen table turns to the lens and says: "Three prompts, ten minutes, one finished clip." Natural window light, shallow focus. No subtitles.

Cases where it holds: one speaker, one line, one location, under 10 seconds. Cases where it breaks: long monologues, two speakers trading lines, names and brand words, anything where the same face has to survive five separate generations.

That last one matters more than it seems. Consistency across clips is the real tax of text-only video. If your series needs the same presenter for 40 modules, a prompt is the wrong tool for the face. Generate the scenes and B-roll with text to video, then let an avatar pipeline carry the speech. Script in, avatar out, scenes cut around it.

Which prompt mistakes should you skip entirely?

Opinion time. These are the habits worth dropping today.

Where it holds: short clips, simple motion, concrete nouns. Where it coils: crowds, hands doing fine tasks, readable text, physics-heavy water and cloth in the same frame.

How do you adapt one prompt across models?

The skeleton travels. The vocabulary does not travel perfectly. Three adjustments are worth knowing before you burn credits.

Runway's guidance is specific here: when you start from an image, describe the motion, not the image. Repeating the image's contents in detail can reduce movement. Kling's guide, for its part, pushes you to think in shots rather than clips, which is the same instinct as the one-action rule above.

Where do you run these prompts, and how do you iterate?

Pick the model for the job, not the logo. Kling and Veo handle camera language and, depending on the version, native audio. Runway is strong on iteration and editing tools. Higgsfield bundles several video models in one workspace, which helps when you want to test one prompt against more than one engine without five logins.

Run the same skeleton prompt on two models. Change one element only: the camera move, the light, or the verb. Keep a log of what each change did. After ten runs you have a house style, and that house style is worth more than any list of examples, including this one.

Laptop with two video timelines beside a studio microphone and ring light on a creator desk

What would you actually do this week with these prompts?

Take one script you already have. Break it into six beats. Write one prompt per beat using the skeleton: camera, scene, one action, details. Generate three takes each, keep the best, and cut them to your voiceover. If the piece needs a face that talks, bring in an avatar for those beats instead of fighting the model for consistency.

You will have a rough cut in an afternoon. The prompts that survive go into a shared doc with their shot type. The next video starts from that doc, not from a blank box.

Frequently asked questions

What is a good text to video prompt?
A good prompt names the camera move, the scene, one visible action and the light or mood, in that order. It describes physical motion the model can render, not abstract feelings, and fits in a single 5 to 10 second beat.
How long should a text to video prompt be?
Long enough to fix the shot, short enough to stay clear. Two to four sentences usually works. Start simple, confirm the motion works, then add one detail at a time.
Should I use negative prompts in AI video?
Mostly no. Runway's guidance favours positive phrasing, because naming what you want gone can put it in the frame. Write 'smooth, stable camera movement' instead of 'no shake'.
Can text to video generate speech and lip sync?
Some current models generate synchronized audio and lip movement for short lines. Results hold for one speaker and one short line. For long scripts or a recurring presenter, an avatar pipeline is more consistent.
Why does my AI video look different every time I run the same prompt?
Generation is stochastic, so each run varies. Lock what you can: use specific nouns, one camera move, and a reference image where the model supports it. Then generate several takes and keep the best.
Which AI video model is best for prompt testing?
There is no single winner. Kling, Veo and Runway each interpret camera language a little differently. Run the same skeleton prompt on two models and compare, changing one element at a time.