# AI Video Prompts: The 7-Element Formula That Works

URL: https://polymorf.me/journal/ai-video-prompts-7-element-formula
Type: blog
Locale: en
Published: 2026-08-15
Updated: 2026-08-21

---

> How to write AI video prompts that reduce unusable clips, which camera terms models actually parse, and the cost of bad prompting hygiene at scale.

Most creators write AI video prompts the same way they type a Google search query. One sentence. A handful of nouns. Hope for the best.

After running 40 training modules through three different pipelines in six months, I can tell you: the gap between a 12-second usable clip and a 12-second unusable one comes down to how precisely you described motion, light, and mood. Nothing else.

![Hands typing on a mechanical keyboard in a dark studio, monitor glowing blue with video thumbnails grid](https://fdzlnqpwsaniezitwiuw.supabase.co/storage/v1/object/public/cms-media/polymorf/2026-08/f4e2e8-inline1.webp)

## Why Your Clips Look Nothing Like What You Imagined

The model is not ignoring you. It is interpreting what you gave it.

"A person walking through a city at night" gives the model five degrees of freedom: which person, which city, which direction, which pace, which light. What you get back is the model's best guess on all five. On a good day, one of those guesses aligns with your vision.

Strong AI video prompts remove those degrees of freedom one by one. You do not need more words. You need more precision.

Here is the difference:

**Weak:** "A businessman walking in a city at night"

**Strong:** "A man in a charcoal overcoat, early 40s, walking briskly left-to-right through a rain-wet Tokyo street at 2 AM, sodium-vapor streetlights above, medium tracking shot, noir mood, slight film grain"

Same subject. The second prompt produces a usable clip on the first try roughly 70% of the time. The first, around 20%. That gap does not shrink over time unless you change how you write prompts. The model does not learn your preferences. It only has what you give it.

For creators building a library of 40 to 100 clips per project, that 50-point gap in first-take success rate translates to hours. At scale, it becomes a staffing decision.

## Avatar Prompts vs. Generative Video Prompts: Two Different Languages

This is the distinction most guides miss entirely, and it will save you hours.

When you are generating **b-roll or cinematic footage** with tools like Kling AI, Higgsfield, or Runway, you are directing a scene. The model needs to understand what is in frame, how it moves, and what it looks like.

When you are generating **avatar talking-head video** with tools like Polymorf, HeyGen, or Synthesia, you are directing a performer. The model needs to understand delivery pacing, avatar mood, and script structure. Not visual scene composition.

Prompting structure for generative b-roll:

`[subject] + [action] + [setting] + [camera] + [lighting] + [style] + [motion quality]`Prompting structure for avatar video (script level):

`[avatar tone] + [pacing cues] + [pause markers] + [emphasis notes]`Trying to write a Kling prompt when you are generating an avatar clip is like giving a cinematographer stage directions. Confusing at best. Counterproductive at worst.

One clip type. Two completely different prompting languages. Knowing which one you are working with before you open the text field saves more time than any formatting trick.

## The 7-Element Stack That Works Across Generative Tools

For b-roll and generative video, layer these seven elements in order. Skip one and you hand control back to the model.

**1. Subject:** Who or what is the focus. Be specific: age, clothing, posture. Not "a woman" but "a woman in her late 30s, navy blazer, standing still, facing the camera".

**2. Action:** What is actually moving. Be kinetic: "slowly rotating," "walking briskly right-to-left," "steam rising from a mug on a marble countertop."

**3. Setting:** Where it happens. Include background detail. "Modern open-plan office with floor-to-ceiling glass, 10 AM light, city skyline behind."

**4. Camera:** Shot type and movement. "Medium close-up, slow push-in." "Wide establishing shot, static." Specific enough to leave no room for interpretation.

**5. Lighting:** Quality, direction, color temperature. "Overcast diffuse light from the left." "Hard directional neon cyan from above." Models parse lighting descriptors accurately when you are precise.

**6. Style:** Cinematic reference or texture. "Kodak Vision3 film stock, slight halation." "Clean corporate look, no grain." "Moody Scandinavian winter, desaturated palette."

**7. Motion quality:** Temporal guidance. Words like "steady," "smooth transition," "no flicker," "consistent movement" dramatically reduce visual artifacts on models like Kling and Higgsfield.

A full 7-element prompt looks like this: "A product designer in her early 30s, grey turtleneck, placing a ceramic mug on a white desk with intention. Modern minimal studio, diffuse window light from the left. Medium close-up, slow push-in. Soft morning light, slightly warm. Clean editorial style, no grain. Steady movement, no flicker."

That prompt generates a broadcast-ready clip on the first or second attempt, consistently. Missing elements 4, 5, or 7 accounts for most of the visual drift and flicker problems creators report after their first week of generation.

![Overhead flat-lay of a filmmaker notebook with shot diagrams next to a camera on a dark minimal desk](https://fdzlnqpwsaniezitwiuw.supabase.co/storage/v1/object/public/cms-media/polymorf/2026-08/8f0940-inline2.webp)

## Camera Movement Terms That Actually Get Parsed

Not all direction language is equal. Some terms are parsed consistently across models. Others get ignored.

**Reliable terms tested across Kling, Higgsfield, and Runway:**

- 
**`static shot`**: Camera stays fixed, no movement at all

- 
**`slow push-in`**: Subtle dolly toward subject, increasing intimacy

- 
**`wide establishing shot`**: Wide angle capturing the full scene

- 
**`medium close-up`**: Chest to head, subject centered in frame

- 
**`tracking shot left-to-right`**: Camera follows horizontal movement

- 
**`slow pan`**: Camera rotates horizontally across the scene

**Terms that frequently get misread:**

- 
"dolly zoom" produces erratic results on most current models

- 
"Dutch angle" is inconsistently applied

- 
"over-the-shoulder" causes the model to center the subject instead

Use the reliable set. If a shot type is not in that table, test it on a short, cheap clip before committing it to a full sequence.

## When Prompt Length Works Against You

Longer is not always better. Each tool has a practical ceiling where comprehension peaks.

For Kling AI: 60 to 100 words is the sweet spot. Beyond that, the model starts deprioritizing elements in the middle of the prompt.

For Higgsfield: similar ceiling, but the model handles parenthetical style notes better. "(subtle, not overdone)" parsed correctly in our tests on six consecutive generations.

For avatar video tools: your script is the prompt. Delivery cues go into pause markers and emphasis notes. One sentence, one clear instruction. A 300-word block does not give you more control. It dilutes it.

The signal that your prompt is too long: you get a clip that nails three elements and ignores two others. Cut the lowest-priority detail first.

## Iterating Without Burning Your Credit Balance

At 20 clips per week with a 40% unusable rate, you are wasting 8 credits weekly. At $0.05 to $0.20 per generation, that is $16 to $80 per month on bad prompts alone. For a team running 100 clips per week, it becomes a real budget line item.

One system that cuts waste: the single-variable test.

Generate a baseline clip with a complete 7-element prompt. Note the result. Change exactly one element and regenerate. Camera movement, lighting, style reference: one variable at a time.

After 4 to 5 iterations, you know which elements your tool weights most heavily. That knowledge transfers to every future clip in your batch.

Rewriting the whole prompt and regenerating tells you something changed but not what. You end up iterating indefinitely.

Save your winning prompt templates. I keep 12 variations organized by setting type: indoor corporate, outdoor urban, product close-up. Reuse the tested skeleton for every new batch.

## What Cuts More Time Than Any Individual Prompt Tweak

The highest-leverage move is not a better formula. It is eliminating the review-and-reject cycle before generation starts.

Two practices that compress production time more than any prompt technique:

**Pre-generation storyboarding:** Write a one-line shot description for each clip before generating anything. If you cannot describe a shot in one sentence, you cannot prompt it precisely either. The storyboarding step forces clarity before you spend a single credit.

**Batch generation with variant testing:** Generate 3 variants of your most uncertain shots simultaneously. Select the best, discard the rest. Faster than sequential iteration and surfaces tool behavior patterns quickly.

For a solo creator producing 2 to 3 videos per week, these two practices reduce generation time by roughly 30 to 40 minutes per project. At 3 projects per week, that is nearly 2 hours of recovered pipeline time.

Production scale. Not studio overhead.

![Organized minimal workspace with tablet showing video production dashboard and color-coded notes](https://fdzlnqpwsaniezitwiuw.supabase.co/storage/v1/object/public/cms-media/polymorf/2026-08/52f273-inline3.webp)

## When the Prompt Is Right but the Clip Still Misses

Sometimes the prompt is correctly structured and the clip still misses. That is not a prompt problem. It is a model alignment problem.

Three situations and what to do:

**The model ignores your camera instruction:** Repeat it twice. Once at the start, once at the end. "Medium close-up, slow push-in. Subject: [description]. Style: [style]. Shot: medium close-up with slow push-in." Redundancy helps on most current models.

**The clip has visual artifacts or flickering:** Add temporal stability language explicitly. "Consistent motion, no flickering, smooth transitions throughout, steady movement." These words are parsed as quality constraints by most video generation architectures.

**The avatar's delivery does not match the script tone:** Break long script sections into shorter paragraphs. Most avatar tools process delivery per-sentence, not per-paragraph. Shorter units give you tighter control over pacing and emphasis. A paragraph that mixes instruction, data, and a call-to-action will produce flattened delivery every time.

When all else fails on a cinematic clip: generate at 5 seconds instead of 10. Short clips have fewer compounding variables. Get the look right on the short version, then extend or loop.

The script is there. The system does the rest, once you have given it something precise enough to work with.

## FAQ

### What is an AI video prompt?

An AI video prompt is the text description you give a video generation model to specify what appears in the clip, how it moves, the lighting, camera angles, and visual style. Depending on the tool, it can range from a single sentence to a structured 7-element block covering subject, action, setting, camera, lighting, style, and motion quality.

### How long should an AI video prompt be?

For generative b-roll tools like Kling AI and Higgsfield, 60 to 100 words is the practical sweet spot. Beyond that, models tend to deprioritize elements in the middle of the prompt. For avatar video tools, your script is the prompt and delivery pacing comes from pause markers and emphasis cues, not raw prompt length.

### What is the difference between a generative video prompt and an avatar video prompt?

Generative video prompts (Kling, Higgsfield, Runway) describe a visual scene: subject, action, setting, camera, lighting, style, and motion quality. Avatar video prompts (Polymorf, HeyGen, Synthesia) focus on script delivery: tone, pacing, pause markers, and emphasis. Applying generative video prompting logic to avatar tools produces poor results because the models are designed for completely different inputs.

### Which camera terms get consistently parsed by AI video models?

Terms that reliably work across Kling AI, Higgsfield, and Runway include: static shot, slow push-in, wide establishing shot, medium close-up, and tracking shot left-to-right. Terms like dolly zoom, Dutch angle, and over-the-shoulder are frequently misread or ignored. When in doubt, test an unfamiliar term on a short, cheap clip before using it in a full production batch.

### How do I reduce unusable clips without changing my creative vision?

Use the 7-element prompt structure and test one variable at a time rather than rewriting entire prompts from scratch. Save your winning templates organized by setting type (indoor, outdoor, product). Batch generate 3 variants of your most uncertain shots simultaneously. Pre-generation storyboarding forces prompt clarity before you spend a single credit.

### Why does the same prompt produce different results across AI video tools?

Each model weights prompt elements differently based on its training data and architecture. Kling AI parses camera movement reliably. Higgsfield handles parenthetical style notes well. Building a baseline understanding of how each tool interprets your key terms saves significant credit spend and prevents endless iteration on prompts that are technically fine but mismatched to the tool.