How to Animate a Photo with AI: Image-to-Video Step by Step
Photos that move
Animating a still image means running it through a video model in image-to-video mode: you give one picture plus a text description of the motion, and you get a clip a few seconds long. Your photo becomes the first frame, and the model builds everything after it.
Below, step by step: which models work in this mode, how to prepare the source image, what lines a motion prompt is built from, and what to do about the usual defects.
What "animating a photo" actually does
The model does not move the pixels of your photo. It takes the image as a condition and generates a new sequence of frames, trying to match the first one to your source and then develop the picture coherently. That is why the further you get from frame one, the more things drift: hair, fabric patterns, the background behind a person.
Three consequences. Anything not visible in the photo (the second ear, teeth, a hand in a pocket) gets invented. The stronger the motion you ask for, the faster the likeness slips away. And motion never improves image quality, it smears it.
There is a fork at the very start: animate an existing photo, or first generate the source image exactly as you want the frame to look. The second route gives you control over composition and a clean source with no noise and no compression artifacts, which is half the result.
What works well and what still fails
Small, predictable motion works well: a slow push in or pull out, a light pan across a landscape, wind in hair, leaves and fabric, steam over a cup, water, fire, clouds, highlights. Portraits work too, as long as the facial motion stays minimal: breathing, one blink, a few degrees of head turn.
It fails where precision matters: hands and fingers, any text in frame, full-body walking or running, changing pose, speech with matching lip movement, handshakes and passing objects.
The test is simple. If the object is fully visible in the source and its motion fits into one verb, your odds are good. If the motion requires new geometry to appear, expect several retries.
Image-to-video models and how they differ
All the models below work in image-to-video mode. The differences are mostly duration, resolution and the character of the motion. The notes come from the model descriptions themselves. I am not quoting prices here: the credit cost of a run is shown next to each model in the interface and it changes.
| Model | What it states | Best for |
|---|---|---|
| Kling 2.1 Standard (image-to-video) | 5 and 10 second clips | basic photo animation |
| Kling 3.0 (image-to-video) | camera movement, multi-shot, 4K mode | portraits and camera work |
| Veo 3.1 (image-to-video) | animation from a reference image, 1080p option | clean, restrained motion |
| Seedance 2.0 Mini | fast and cheap mode | drafts and idea sweeps |
| Seedance v1 Lite (image-to-video) | 480p/720p/1080p, 5 and 10 seconds | inexpensive prompt tests |
| Wan 2.6 (image-to-video) | 720p/1080p, 5, 10 and 15 seconds | shots longer than 10 seconds |
| HappyHorse Image-to-Video | reference image required | animation kept close to the source |
In practice: run the prompt on a cheap, fast model until the motion is what you wanted, then repeat the same prompt on a more expensive model at the resolution you need.
[[МОДЕЛИ: Kling 3.0 (image-to-video) | The model list scrolled to video models in image-to-video mode]]
Step by step: from uploaded photo to finished clip
- Prepare the source: the crop you want, no watermarks, and if the face is the point, it should not be smaller than roughly a quarter of the frame height.
- Decide the aspect ratio first and crop the photo to it yourself, otherwise the edges get painted in, often badly.
- Upload the image to a video model that takes an image as input.
- Write the motion prompt using the skeleton from the next section: one photo, one set of linked movements, no three-scene story.
- Set the shortest duration, usually 5 seconds. The test is cheaper and shows the same defects as a ten second clip.
- Watch the result twice: once straight through, once frame by frame over the last second. Defects almost always pile up at the end.
- Change one thing per pass: prompt first, then duration, then model.
[[СКРИНШОТ: /dashboard?type=video&q=animate photo | The input field with an uploaded photo and a motion description, right before the video model runs]]
Writing the motion prompt: camera, subject, light
A prompt for animating a photo does not describe the picture, the picture already exists. It describes only what should change. A five-line skeleton:
Camera: [one move] - slow push in / slow pull out /
light pan left / static camera
Subject: [what moves] - hair and fabric in a light breeze,
the person breathes, blinks, turns the head slightly
Environment: [background and details] - steam rising from the cup,
leaves swaying in the background, highlights on water
Light and style: light, colors and grain as in the source, no recoloring
Limits: 5 seconds, no sharp head turns, no pose change,
no new people or objects, no textFilled in for a portrait:
Slow camera push in, about 10 percent. Hair and collar move
slightly in the wind, the person breathes, blinks once, turns
the head a few degrees and turns back. The background stays the
same, light and colors as in the source. 5 seconds. No face
distortion, no sharp movement, no change of expression into a smile.What matters here: one camera move instead of "push in while orbiting"; verbs ("blinks", "sways") instead of adjectives ("dynamic", "cinematic"); a limits line that cuts pose changes, invented extras and recoloring; numbers, because "push in 10 percent" is more stable than "move the camera closer".
If the motion is hard to imagine, ask a text model to describe it: upload the photo, ask it to list what could physically move in this frame, and build the Subject and Environment lines from the answer.
[[СХЕМА: five lines of a motion prompt - camera, subject, environment, light, limits - and a filled-in portrait example | What a photo animation request is made of]]
Source image requirements: resolution, crop, noise
The model amplifies what is already in the photo, defects included. Rough targets:
- Resolution. Short side no less than 1000 pixels.
- Face size. If the face is a hundred pixels across, the features will drift in motion almost every time. Crop tighter.
- Noise and compression artifacts. A scan of an old print, a re-saved JPEG, a screenshot from a messenger: upscale first, animate second, otherwise the grain starts to boil.
- Crop. Leave some air around the subject if you plan a push in.
- Remove in advance. Watermarks, stray captions, someone else's hand at the edge of the frame, a cable across a face.
- Format. JPEG or PNG with no transparency. A transparent background gets filled in however the model sees fit.
A five second check: open the photo at 100 percent and look at the eyes and fingers. If it is already mush there, video will not fix it.
Common defects and what to do about them
The face loses likeness. The model redraws the head on every turn. Fix: a tighter portrait on input, "static camera", a turn of no more than a few degrees, 5 seconds instead of 10.
The eyes go their own way. Ask for one blink per clip and write "gaze directed at the camera". "Lively eyes" gives you a squint.
Hands. Safest is not to ask for hand motion at all, and if they get in the way, crop so the hands fall outside the frame.
Text warps. Lettering does not survive generative video. Either crop the text out, or overlay it on the finished clip in an editor.
Flicker and boiling texture. Usually a noisy source or motion that is too fast. Clean the input, slow the camera down.
Background morphing and a jump at the end. The line "background stays unchanged, no new objects appear" plus a shorter duration both help.
If the same defect survives three attempts, the problem is the source image or the model, not the prompt.
After the render: upscale, sound, stitching
Upscale. If you generated at 720p to save credits, raise quality with a separate model: Topaz Video Upscale states up to 4x, and for stills Topaz Image Upscale goes up to 8x. Order matters: upscale the photo before the video, upscale the clip after.
Sound. Wind, footsteps, a creaking door come from ElevenLabs Sound Effects by text description, a music bed from the Suno models, and you mix them in a video editor.
Stitching a longer scene. Export the last frame of the clip as an image, feed it as the source for a new generation, continue the same motion in the prompt, and cut the pieces together with no transition. Quality usually drops by the third or fourth link, because each new frame is a copy of a copy. For anything longer than half a minute it is more honest to generate several independent shots and edit them as different angles.
How to animate a photo on Moleculs.ai and what it costs in credits
The order is the same: find a video model in the catalog that takes an image as input, upload the photo, write the prompt using the five-line skeleton, pick a duration and run it.
Runs are paid for in credits from your subscription. Each model has its own cost per run and it is shown next to the model, so you see the price before you start. Video generations cost more than text ones, so the free quota of 5 credits will not cover a clip. Paid quotas refresh every billing period, and the current numbers are on the pricing page.
A sane spending habit: run drafts on fast, cheap video models at 5 seconds until the motion is right, and only repeat the final version on an expensive model at full resolution.
A first prompt worth trying: "Animate this photo: slow camera push in, light movement of hair and fabric in the wind, the person blinks and turns the head slightly, light and colors as in the source, 5 seconds, no face distortion." Attach your own photo and compare two models on it. That is the fastest way to learn which one behaves more carefully on your kind of frames.
Related reading

Corporate access to AI models
Invoice for legal entities, centralised payment, priority support
- Access to ChatGPT, Gemini, Grok, Claude and DeepSeek
- Prompt library and shared access inside the team