Blog
September 21, 2026

How to Generate Video with AI: A Step-by-Step Guide

To get a short clip out of a neural network you need three things: a scene described in one paragraph, a choice of mode (from text or from an image), and the understanding that one generation gives you a few seconds of footage, not a finished clip. Everything else is prompt detail and patience for retries.

Below is a step-by-step breakdown: what to pick, what to type in the prompt field, how to animate your own photo, how video models differ, and why the result drifts away from the picture in your head. The order runs from simple to complex: a single generation first, then camera, sound, upscaling and assembling a longer clip.

What you need before the first generation

The minimum kit: an account on Moleculs.ai with access to video models, a stock of credits, and one scene you have already worked out in your head. Video generation costs noticeably more than text or images, so blind "let's see what happens" runs are expensive.

Before you open the prompt field, answer four questions:

  • What happens in the frame? One action, not three.
  • Where does the camera sit, and does it move?
  • What is the light and the time of day?
  • Do you need a source frame, or should the model invent the scene?

If your answer to the first question sounds like "the character walks into the room, sits down, opens a laptop and starts talking", the clip will almost certainly fall apart. Five seconds hold one action: steam rising, the camera pushing in, a person turning their head. The rest becomes separate shots.

Decide in advance where you will assemble the result. A video editor is needed even for a simple clip: to join two shots, drop music underneath, trim half a second off the tail.

Two modes: text-to-video and image-to-video

Text-to-video: you describe the scene in words and the model invents both the composition and the motion. Fast, but unpredictable. Faces, object proportions and spatial logic go wrong here most often.

Image-to-video: a finished frame becomes the first frame of the clip, and the model is only responsible for movement. More control, less waste. The price of that control is that you need a source image.

The workflow that fits most tasks: push the frame with an image model until you think "yes, that's it", then send it into image-to-video with a description of the motion. The difference in the share of usable generations shows up immediately.

Text-to-video in 6 steps

  1. Find a video model in the catalog. The launch cost in credits is shown next to each one, and that is your first filter.
  2. Pick the model for the job: a cheap one to test the idea, a heavier one for the final take.
  3. Set the parameters the model lets you change: duration, resolution, aspect ratio. Proportions belong here, not in the prompt text.
  4. Write the prompt using the structure from the prompt section below: scene, subject, action, camera, light, style, constraints.
  5. Run the generation. Video takes longer to compute than an image, which is normal.
  6. Watch the clip twice: all the way through, then paused on the first and last frames. Extra fingers, warped objects and camera jerks show up at the edges.

After that the decision is simple: a mistake in the story or composition means fixing the prompt and regenerating, a softness problem means upscaling.

Animating a photo: image-to-video step by step

  1. Prepare the frame, ideally in the same aspect ratio as the future video, otherwise the model will crop the edges.
  2. If you have no source image, you can make the first frame with an image model and animate it afterwards.
  3. Choose a model in image-to-video mode and upload the picture.
  4. In the prompt, describe motion only. The model already sees what is in the photo.
  5. Rule out the extras: "the camera does not change angle", "the background stays still", "no new objects in frame".
  6. Check the first frame of the result: it should match your source. If the model rewrote it, dial the motion down in the description.

An example prompt for animating a portrait photo:

Animate this frame: the person slowly turns their head toward the camera,
slight hair movement from a breeze, eyes blink once.
Camera static. Background unchanged. No new objects.
Duration 5 seconds.

The typical mistake is asking for too much movement. "The character stands up and walks" in five seconds from a static photo almost always gives you melting anatomy.

Video models and how they differ

The choice comes down to three things: supported mode, stated duration and resolution. Below is what the models claim about themselves. The catalog changes, so check the description next to the model.

ModelModesWhat is stated
Veo 3.1text and image1080P option
Veo Omnitext and imagemultimodal input, up to 4K
Kling 3.0text and imagemulti-shot, element references, 4K mode
Kling 2.1 Standardimage5 or 10 seconds
Seedance 2.0text and imagemultimodal references
Seedance 2.0 Minitext and imagefast and cheap option
Seedance v1 Liteimage480p/720p/1080p, 5 or 10 seconds
Hailuo 02 Standardimage512P/768P, 6 or 10 seconds
Wan 2.6image720p/1080p, 5, 10 or 15 seconds
HappyHorsetext and imageup to 1080P, up to 15 seconds; reference image required in image mode
Grok Imagine Videotextvideo from a text description
Grok Imagine (image-to-video)imageanimates an uploaded frame

For drafts, take a cheaper and shorter model; for the final take, the one that states the resolution you need. If you need a change of shot inside a single fragment, look at models that state multi-shot. We do not build these models and we do not promise that behaviour in our interface matches the vendor's own app: the wrapper and the settings can differ.

Writing a video prompt: structure and a worked example

A video prompt differs from an image prompt by one block: it has time in it. On top of what is in the frame, you have to say what changes over those seconds.

A working skeleton you can copy and fill in:

Duration: 10 seconds
Scene: [where it happens, time of day]
Subject: [who or what is in frame, shot size]
Action: [one change across the whole fragment]
Camera: [static / push in / pull out / lateral dolly, speed]
Light: [source, direction, hardness]
Style: [realism / film / 3D render, palette]
Constraints: [no text, no new objects, no hard cuts]

A filled-in example:

Duration: 10 seconds
Scene: kitchen by the window, early morning
Subject: a cup of coffee on a wooden table, close-up
Action: steam rises from the coffee and slowly curls
Camera: slow push in, no jerks
Light: morning sun from the side, soft shadows
Style: realism, warm colors, shallow depth of field
Constraints: no text or captions, no hands in frame

Why this works:

  • one action (the steam) instead of three;
  • one camera move instead of "push in, then orbit";
  • light described by source and direction, not by the word "beautiful";
  • constraints spelled out: captions and watermarks are things models love to add on their own.

The same prompt written as a single line works too, as long as you keep the order of the blocks.

Camera moves and cuts: what you can actually control

Simple camera instructions land reliably: static camera, slow push in, pull out, lateral dolly, pan left or right, handheld with light shake. Write them one at a time and add speed: "very slow push in" and "fast push in" give different results.

Complex combinations behave badly, things like "a 180 degree orbit around the object, then cut to a close-up". On a short fragment the model either picks one of them or breaks the geometry of the scene.

Changing shots inside one generation is a separate story. Some models state multi-shot, meaning several shots in one clip. If that is not available, you cut by hand: two fragments and an edit point in your editor.

A useful habit: put "no hard cuts and no change of angle" in the constraints when you need one continuous shot. Without it, models sometimes insert a transition on their own initiative.

Sound: effects, music and voice over the clip

Some video models return silent video, some return sound, and the description does not always let you predict which. It is more practical to assume you will build the audio yourself.

What it is built from:

  • noises and effects (footsteps, rain, city hum) generated from a text description by a separate sound effects model;
  • a music track from the Suno family of models, from short beds to tracks with vocals;
  • voice, either recorded or prepared in advance as a file you place under the video.

Then everything comes together in a video editor: a video track, an effects track, a music track. There is no editing timeline inside Moleculs.

The order matters: final video fragment first, then effects tied to specific actions in frame, music last. Otherwise you end up fitting the video to the rhythm of the track, and regenerating video is the expensive part.

Upscaling and building a clip longer than one generation

A soft or low-resolution result is fixed by upscaling: Topaz Video Upscale goes up to 4x. It raises resolution and detail but does not rework the content, so extra fingers and a jittery camera only become more visible.

A clip longer than one generation is assembled from fragments:

  1. Break the idea into 5-10 second fragments, one action each.
  2. Generate the first fragment.
  3. Take its last frame as the source for the image-to-video pass on the next fragment.
  4. Keep the wording for style, light and palette identical across all prompts.
  5. Stitch the fragments in an editor, with short transitions if needed.

A shot list is easy to prepare as text: you can ask a text model to break the prompt down shot by shot and get a list of fragments with the action and camera for each. Then you run that list through a video model.

The seam between fragments will still be visible in small details: a slightly different tint, a slightly different texture. A single color grade across the whole clip at the end helps.

Why the video comes out wrong: 7 common mistakes

  1. Too many actions in one fragment. Five seconds hold one change.
  2. Two camera moves at once. Push in plus orbit in one prompt gives mush.
  3. No constraints. Without "no text or captions", the model may write a caption into the frame.
  4. The prompt describes a still image. No verbs of motion means a nearly frozen shot.
  5. Aspect ratio set in words instead of a parameter. "Vertical video" in the prompt works worse than the setting.
  6. Text-to-video where a specific object matters. A product, a face or a logo is safer fed in as an image.
  7. Rewriting the whole prompt after one failure. Change one block at a time.

The pass criterion: the first and last frames look fine on pause, motion runs in one direction without jerks, and there is nothing in the frame you did not ask for.

How credits are spent on video

Inside a subscription you spend credits, and every model has its own launch cost, shown next to it in the catalog. Video generations cost more than text requests, so savings come from the order in which you work, not from the plan you pick. Quotas and what each plan includes are on the pricing page.

What actually reduces spend:

  • test the idea on a fast, cheap model and do the final take on a heavy one;
  • get the first frame right with an image model first, where a retry costs less;
  • generate short durations while you are still tuning the prompt;
  • do not fire off repeats without editing the prompt: the same request gives a similar result.

The credit quota renews every paid period. The free plan gives 5 credits, which is enough to get acquainted with text models but not with video. Upscaling and sound generation also consume credits, so plan them for the end, once the fragment is approved.

Corporate access to AI models

Invoice for legal entities, centralised payment, priority support

  • Access to ChatGPT, Gemini, Grok, Claude and DeepSeek
  • Prompt library and shared access inside the team
Request an invoice