Production workflow guide · Updated August 24, 2026
How to make a photo talk with your own audio
A talking-photo model combines the identity and appearance in a portrait with timing from a speech recording. It predicts facial motion; it does not recover what the person really said, felt, or looked like while speaking.
What you need
One portrait
JPG, PNG, or WebP, up to 10MB. A centered, unobstructed face is the safer starting point.
One speech file
MP3, WAV, or M4A, up to 10MB. TalkingPhoto currently uses uploaded audio rather than built-in text-to-speech.
Step-by-step workflow
- 1
Choose one clear portrait
Use an image with one visible, front-facing face. Avoid heavy blur, deep shadows, covered eyes or mouth, and very small faces.
- 2
Prepare a speech recording
Use an MP3, WAV, or M4A file. Clean speech with limited background noise gives the model a clearer timing signal.
- 3
Upload and validate
TalkingPhoto checks file ownership, type, size, and image readability. Use a clearly visible face because the current upload step does not promise automatic face detection.
- 4
Submit the generation job
A credit is reserved, then an asynchronous task calls the pinned SadTalker model. The job may remain queued or processing while the provider works.
- 5
Review before publishing
Watch the entire MP4 for lip, teeth, eye, skin, and head-motion artifacts. Regenerate with a clearer portrait or cleaner audio if necessary.
Why portrait choice matters
The model has to infer depth and motion from a single still image. A strong side profile, cropped chin, covered mouth, tiny face, or uneven lighting removes useful evidence. Old photos can work, but age alone is not the deciding factor: face visibility and image clarity matter more.
Limitations and responsible use
Generation time varies with provider capacity, input length, retries, and network conditions. A successful job can still contain visual artifacts. Only upload photos and voices you own or are authorized to use, obtain consent where required, and disclose synthetic media when viewers could reasonably mistake it for a real recording.
Do not use a talking-photo result for deceptive impersonation, fraud, harassment, or misinformation. A memorial or family-history project deserves particular care because the generated speech is an artistic reconstruction, not historical evidence.
Ready to test your own inputs?
Review the supported formats and limitations first. When generation is available, a newly authenticated account may use one private, view-only preview of up to 5 seconds; downloadable results use purchased credits. Completion depends on the configured AI provider being available.