Input guide
Prepare a portrait and speech recording that an animation model can actually use
Most avoidable failures begin before the upload. This checklist improves the source material without promising a result the external model may not deliver.
Confirm permission before editing
Use a portrait and recording you own or have permission to animate. Write down the intended audience and obtain explicit consent when a real person is identifiable.
Crop around one clear face
Use a front-facing or slightly angled portrait. Keep the eyes, nose, mouth, jaw, and some space around the head visible. Avoid a face that occupies only a tiny part of a group photo.
Use even light and a neutral mouth
Strong shadows, hands, microphones, sunglasses, and hair across the mouth remove information the animation model needs. A closed or relaxed mouth is a safer starting frame.
Record clean, dry speech
Record in a quiet room, keep a consistent distance from the microphone, and remove music or heavy reverb. Export MP3, WAV, or M4A; TalkingPhoto validates the actual file rather than trusting its extension.
Trim silence and check duration
Remove long lead-in and trailing silence. TalkingPhoto accepts up to 60 seconds and bills each submitted clip in started 10-second blocks, so 20.1 seconds uses 3 credits.
Listen once before submitting
Check for clipped words, sudden volume jumps, and private information. The AI task cannot repair a sentence you did not intend to publish.
TalkingPhoto upload boundaries
- Portrait
- JPG, PNG, or WebP · 10MB maximum · normalized to JPEG after validation
- Speech
- MP3, WAV, or M4A · 10MB maximum · 60 seconds maximum
- Free preview
- One eligible view-only job · up to 5 seconds · retained for 24 hours
- Paid output
- One credit per started 10 seconds · private download link · 30-day retention