Short answer: Turning photos into video is two different jobs. Making a slideshow or clip from several photos is editing; your phone's gallery or editing app is enough. Animating a single photo with AI is image to video: the photo becomes the first frame, you describe the camera and subject motion, and the model generates a clip of 4 to 10 seconds.
A short memory video from a wedding album, a rotating showcase shot from a product photo, or a smile appearing on your grandfather's old portrait: all three get typed into a search box as "photo to video", but they are not the same job. Start down the wrong route and you either spend tokens for nothing or never get the motion you expected. This article separates the two jobs and explains, for animation, how to choose the photo, how to write the prompt and what each option costs.
What does turning a photo into a video mean?
This search covers two different jobs. The first is editing: several photos are put in order, with transitions, slow zooms, music and text between them; nothing inside the photos moves. The second is image to video: a single photo becomes the first frame and an AI model generates the hair moving, the head turning or the camera moving forward.
| Slideshow video from photos (editing) | Animating a single photo (image to video) | |
|---|---|---|
| Input | Several photos | One photo, the first frame of the clip |
| Motion inside the photo | None, only the framing moves | Yes, the subject and the camera move |
| What appears on screen | Only your own photos | Frames that are not in the photo are predicted by the model |
| Length | Depends on the number of photos, can run for minutes | 4 to 10 seconds per clip |
| Tool | Your phone's gallery or editing app | An AI studio that supports image to video |
| Cost | Usually free | Tokens per clip |
| Best for | Memory videos, event recaps, property tours | Product shots, bringing portraits to life, scene openings |
The two can be combined: animating several photos one by one and then lining them up in the edit gives you both the authenticity of the photos and the motion. We explain how image to video differs from text to video in the detailed comparison article.
How do you make a video from photos on your phone?
You do not need AI to make a video from the photos on your phone. The memories or movie feature of your gallery app, or any editing app, puts the photos in order, gives each a few seconds, and adds music and transitions. This route is free, and because it adds nothing to the images, it is the most honest option for memory videos and documentary work.
- Choose the photos in the order of the story; out of three similar shots of the same moment, keep only the best.
- Decide the target ratio up front. Phone cameras usually shoot photos at 4:3 to use the whole sensor; since vertical video is 9:16, the edges get cropped. Keep important faces in the middle of the frame.
- Give each photo a duration that matches the tempo of the music: short for fast music, long for calm music.
- Add a slow zoom or a horizontal pan. This effect, familiar from documentaries, gives a still photo a sense of motion.
- Add text and dates in the editing app; that way they stay legible and special characters display correctly.
How does photo animation with AI work?
In an image to video model your photo is the first frame of the clip and stays fixed. The model generates the following frames from its prediction of how the scene in the photo will move; your prompt steers that prediction. Because the composition comes from the photo, the face and the product stay recognizable, but the motion is limited by the edges of the frame.
This method has a consequence: anything not visible in the photo is made up. If the head turns, the back of the ear is the model's guess; if the camera pulls back, so is the room outside the frame; if the hand opens, so is the palm. The longer the clip, the more the guesses pile up; with many models the face starts to change slowly after five or six seconds. That is why short clips and limited motion give the most reliable results. We explain how we use the same principle in our film pipeline in the keyframe first article.
Which photos work well for animation?
A photo that works well is sharp, well lit, has a single subject and a simple background. The face is seen from the front or slightly from the side, the hands are either fully in frame or not there at all, and there is no text in the frame that needs to be read. Because the model magnifies every flaw in the photo through motion, choosing the source makes more difference than writing the prompt.
- Resolution: The photo should not be smaller than the target video. A 1080p video means 1920 × 1080 pixels; a photo whose short side is well below that goes soft when it is enlarged.
- Ratio: Crop the photo to the video's ratio yourself beforehand. If you leave the cropping to the model, an important detail may end up outside the frame.
- Face angle: Front-on or three-quarter view, eyes open and visible. Turning a face photographed in profile toward the camera means inventing half the face.
- Background: A simple, out-of-focus background is best. Faces and objects in a busy background melt during motion.
- Hands: Hands should either be fully visible or out of frame. Half-cropped fingers are the detail that breaks most often.
- Text, logos, borders: Signs, labels and logos break up when they move. Photos with a border or text added on top are read like posters and reinterpreted.
- Old photos: Remove scratches, stains and tears in a restoration step first; the model may take a stain for part of the scene and set it moving.
How do you write a motion prompt?
A motion prompt is built from three parts: camera movement, subject movement and tempo. Do not describe the subject that is already visible in the photo again; write only what will change. A single sentence such as "The camera slowly pushes in, the woman smiles slightly and turns her head to the left, her hair stirs in the wind, calm tempo" is often enough.
| Weak prompt | Why it is weak | Better prompt |
|---|---|---|
| "Bring the photo to life" | No motion is defined, so the model picks at random | "The camera slowly pushes in, the man blinks and smiles slightly" |
| "A young woman in a yellow dress on the beach" | It describes the subject of the photo again | "Her skirt ripples in the wind, waves break on the shore, the camera is static" |
| "She runs, hugs her friend and dances" | Three actions in one clip | "She takes two steps toward the camera" |
| "Make the product spin, the logo shine and text appear" | Text and logos break up | "The bottle slowly turns a quarter turn on its axis, studio light slides across the surface" |
A few patterns are enough for camera movement: slow push-in, pull-back, slide to the right or left, a gentle tilt up, and a static camera. The camera and the subject moving fast at the same time is the most common cause of breakdown; if you move one, keep the other calm.
Where does photo animation break down?
Breakdown starts where the model has to make up information that is not in the photo. There are five common cases: fast turns, hands, text, crowds and long duration. Preventing them through photo choice and short clips, rather than forcing them with the prompt, means fewer regenerations and fewer tokens.
- Fast turns: A head or body turning through a wide angle forces the model to invent the unseen face; the person can turn into someone who does not look like them.
- Hands and holding objects: A hand lifting a glass can change the number and shape of its fingers.
- Text: Text on a sign or a T-shirt turns into gibberish within a few frames.
- Crowds: In a group photo the faces in the back row blend into each other; the model cannot tell who the subject is.
- Long clips: As the length grows, the face and clothing drift. If you need a long moment, generate two short clips and join them in the edit.
Is it right to animate a photo of a deceased loved one or of someone else?
It is technically possible, but the question of consent and context comes before the technology. Bringing a photo of a deceased relative to life as a memory within the family is different from spreading the same image publicly. Animating someone else's face, or making it speak, without their permission carries both ethical and legal risk.
In Türkiye, Personal Data Protection Law No. 6698 treats a photo of an identified or identifiable person as personal data. Article 50 of the European Union AI Act has also required, since 2 August 2026, that generated images resembling real people, places or events, which could be taken as genuine, be clearly labeled as artificial. This is not legal advice; for commercial or public use, consult a lawyer.
- Even for a family memory, ask the relatives of the person in the photo; not everyone receives an animated face with the same feelings.
- Do not animate photos of children and share them publicly; bear in mind that they cannot give consent.
- State in the video you share that it was animated with AI, and do not present the footage as a real recording.
- Do not use photos of well-known people, your customers or your employees without written permission.
We explain how visible and invisible labels work in generative media in the watermarking and provenance article.
How many tokens does it cost to turn a photo into a video on Yeşilçam?
In Yeşilçam Studio you upload your photo as an asset; the frame you provide becomes the first frame of the clip and is animated at the quality tier you choose. The price depends on the tier and the length, is shown on screen before every generation, and failed generations are refunded automatically. As of October 2026, the image reference line in the price list is 2,626 tokens.
| Tier | Length | Clip (tokens) | Total with image reference | ≈ $ at the Basic rate |
|---|---|---|---|---|
| Draft | 5 s | 122,501 | 125,127 | 0.75 |
| Standard | 5 s | 166,251 | 168,877 | 1.01 |
| Cinema | 6 s | 262,501 | 265,127 | 1.59 |
| Cinema Pro | 6 s | 700,001 | 702,627 | 4.22 |
The dollar column is calculated at the Basic plan's rate: 5 million tokens for $30, so 1 million tokens is about $6. On the Free plan a ×1.3 multiplier applies; a 5-second Draft clip with an image reference comes to 125,127 × 1.3 = 162,665 tokens and fits into the monthly 250,000-token allowance once. Free output carries a watermark, and the Cinema tiers unlock with the Basic plan. On the Pro plan the same operation, at ×0.8, uses 100,102 tokens.
A 30-second memory video from six photos
Animating six photos for five seconds each at the Standard tier and joining them costs 6 × 168,877 = 1,013,262 tokens; on top of that come 5,834 tokens for assembling the film and 202 tokens for the 1080p render. The total is 1,019,298 tokens, about $6.12 at the Basic rate. Lining up the same six photos as a still slideshow in an editing app is free; the difference is the price of the motion inside the photos.
The rule for tiers is simple: first check at Draft whether the motion is right, then produce it at Standard if it works, and keep the Cinema tiers for the one photo where light and motion really matter. You will find the differences in the draft, standard, cinema article and current prices on the pricing page. If you want to animate the photo as an illustration rather than realistically, you can upload it as a character reference in the Animation studio (JPG, PNG or WebP, up to 25 MB); the style you choose, Claymation for example, translates the photo into its own language. If you are moving on to a narrated job with several scenes, the how to make videos with AI guide covers the whole pipeline.
What does photo animation not solve?
Animation creates a moment; it does not record one. The resulting clip shows not how the person in the photo really moved that day, but a motion the model finds plausible. For work with documentary value, such as a family history documentary or news, still photos and editing are the more honest choice.
- It does not sharpen a blurry or very small photo; restoration and upscaling are separate jobs.
- Making the person in the photo speak requires voice, a script and lip sync, and that also requires separate permission from the person or their relatives.
- It does not produce a long scene; clips of 4 to 10 seconds are joined in the edit.
- In a product photo, it does not prove how a fabric drapes or the real texture of a material; details that affect the buying decision need a real shoot.
Frequently asked questions
Is turning a photo into a video free? Making a slideshow video from photos is free with your phone's gallery or editing app. Animating a single photo with AI costs tokens per clip; Yeşilçam's Free plan gives a monthly allowance that covers one 5-second Draft clip, and the output carries a watermark.
How many seconds does an animated photo last? A clip runs 4 to 10 seconds depending on the tier: Draft and Standard 5 or 10, the Cinema tiers 4, 6 or 8 seconds. For a longer video, generating several short clips and joining them in the edit gives a more consistent result than one long clip.
Can old black-and-white photos be animated? Yes, but remove scratches, stains and tears first; the model may take them for part of the scene. Color is a separate decision: if you do not state clearly in the prompt that it should stay black and white, the model may add color, and that can spoil the photo's period feel.
Can I make the person in the photo speak? Technically it is possible with voice, a script and a lip sync step; on Yeşilçam lip sync starts at 29,167 tokens. But making someone say words they never said is the most sensitive use. Do not do it without the explicit permission of the person or their relatives, and state in the video that it is artificial.
In what format does the animated video download? On Yeşilçam the assembled film is rendered in 720p, 1080p or 4K and downloads as MP4, WebP or GIF. MP4 suits delivery and sharing, WebP or GIF in-page previews; choose the ratio up front based on where you will share it.

