Feature
Photo to video avatar: seven rules the portrait has to pass
Most presenter renders fail on the photo, not on the model. The picture was chosen to look good in a grid. A render needs a picture that can move. Those are different jobs. Switching tools does not fix it. Here is what the uploader takes, the seven rules a portrait has to pass, and how Cutroom turns the render into a finished ad.

- Largest photo the uploader takes
- 15 MB
- Per minute of finished render
- 320 CR
- Free credits, no card, one test clip
- 300
What the uploader takes, and the one rule with no workaround
The uploader takes a PNG, JPEG or WebP. Up to 15 MB. Anything larger is refused before it costs a credit.
The audio runs from three seconds to five minutes. Record it on a phone, or generate a read here for 70 credits a minute.
Out comes a lip-synced 9:16 MP4. That face, chest up, saying your words.
Now the rule with no workaround. The person in the photo is you, or somebody who gave you written permission.
Not a photo you found. Not a stock headshot licensed for print. Not a colleague who said yes out loud in a meeting.
Keep the permission on file. Note the date. Note what was agreed.
People leave. The ad keeps running.
Cutroom asks you to confirm this before the first render. Everything after it is craft. This part is not.
Seven rules, and ten minutes by a window beats any photo you own
The model works out how a whole face moves from one frozen instant. Every ambiguity in the still turns into an artefact in the video.
So shoot one on purpose. Ten minutes with a phone near a window beats the best photo in your camera roll.
The photos you already own were framed to look good. They were never framed to move.
Treat it as building a set rather than picking an image. You are choosing the room, the light, the shirt and the expression.
Every script for the next six months gets delivered in that set. Nobody would spend two seconds on that decision at a real shoot.
Shoot five while you are stood there. Different shirt. Different expression. One step further from the wall.
Two of the five will render noticeably better than the other three. You cannot pick which two by looking at the stills.
Render a short test from the best two. The 300 trial credits cover it and no card is needed.
- Face square to the lens. A three-quarter angle photographs better and renders worse.
- Eyes to the lens. Looking off camera means the finished video addresses nobody.
- Even soft light on the face. A hard shadow across one cheek gets baked in, then has to move.
- Mouth closed or barely parted. A wide smile is the most common single cause of a strange result.
- Nothing crossing the face. No hand near the chin, no microphone, no hair over the mouth.
- A plain wall behind. Detail behind a moving head is where artefacts get spotted.
- Shoulders in frame. A tightly cropped head leaves no body for the motion to belong to.
Ten minutes by a window
The portrait shoot that costs nothing and fixes most render complaints.
Stand side on to a window
Daylight, no lamp fighting it
Step away from the wall
Two paces kills the hard shadow
Phone at eye height
Lens level, not tilted up
Mouth closed, eyes to lens
Neutral beats a big smile
Shoot five, render one
300 free credits covers the test
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Five photos that look fine and render badly
A hat brim, a high collar, a scarf near the jaw. A rigid edge close to the moving region reads as detached from the face.
A photo that has already been processed. Beauty filters and skin smoothing strip out the detail the render needs.
Portrait-mode blur is worse again. The blurred background smears the moment the head starts moving.
A low-resolution image, or a screenshot of one. Upscaling before you upload does not put information back. It makes the missing information smoother.
A dramatic expression. Mid-laugh, mid-word, eyebrows up. The video has to travel away from that face and the journey is what people notice.
Glasses with a strong reflection. The reflection stays pinned to the frame while the head moves underneath it.
None of these throw an error. They return a clip that feels wrong for a reason nobody can name.
Reshoot rather than retouch. Retouching caused three of the five.
What each input decides, ranked
Buyers spend the afternoon on the portrait. It is the input with the least room to help.
- The script, read out loudMost of it
- The audio recording qualityThe next slice
- The portraitRules options out
The render is a take. This is where it becomes an ad.
What comes back is one continuous shot of a face. Posting that on its own is why people decide the format looks cheap.
Most presenter tools end at that file. You export it and start again somewhere else.
Cutroom does not end there. Send the MP4 into the editor as a source take. One batch is 100 credits.
It comes back as a finished ad, and you direct it on the transcript.
Highlight a phrase and a clip lands over exactly those words. Delete a line you do not want and the cut rebuilds around it.
Change the pace and the machine re-cuts the whole piece.
At normal pace up to 45 percent of the runtime is covered with footage. The portrait is on screen for barely half the ad.
Captions come from six style packs. Most of your audience meets the words with the sound off.
Music sits under it for 20 credits a track. Export is 20 credits per output minute.
That second step costs less than a third of the render it is finishing.
The set your next six months of scripts get delivered in
One portrait fronts every script where the information is the value.
An explainer. An offer. A spec walkthrough. A policy change. A restock notice.
Nobody wants to film those five times. The render delivers them identically in January and in June.
That is the real advantage. Render five versions of the same offer in an afternoon. Let the numbers pick the winner.
Some scripts need a product held up to the lens. There are no hands in a rendered frame.
Film that part on a phone. Let the presenter carry the words around it.
B-roll is placed on the transcript here, so joining the two is one highlight.
The plain facts to plan around. Output is a 9:16 MP4. Audio caps at five minutes. A filmed source take caps at three.
There is no timeline, by design. You direct on words instead of a scrubber. The decisions stay about the ad.
That second half is what no presenter tool offers, because none of them own the edit after the render.
Questions people ask
- Can I use a photo generated by an image model?
- Yes, and it removes the permission question because nobody is being depicted. Generated portraits render well when they are front on, evenly lit and unprocessed, exactly like a real one.
- Does a higher resolution photo make a better video?
- Up to a point, then it stops mattering. A clean 1024 pixel portrait shot by a window beats a 4000 pixel one lit from one side. Sharpness at the mouth and eyes is what counts.
- How many photos should I make before I commit?
- Five, shot in one session, then render a short test from the best two. The 300 trial credits cover roughly one minute of render, which is two twenty-second tests.
- Is it right for every ad?
- It is built for scripts where the information is the value. If your ad is your own story, film it on a phone and run that take through the same editor for 100 credits.
Ten minutes by a window, five frames, one test render. Then the editor finishes the ad. Other presenter tools hand you the file and stop.