OmniFit logo OmniFit Blog home
Talking-photo workflow

How to make a still image talk with AI tools

If you want a quick talking-photo setup without building a full edit from scratch, prep the character still, generate the voice track, then bring both pieces into HeyGen to make the avatar speak.

📅 Published: May 9, 2026
Image to video Simple tool stack Best for explainers
What this post teaches

A short three-part workflow for turning one static character image into a ready-to-play talking slide. Think of a talking Albert Einstein, a classroom mascot, or a simple presenter-style avatar.

Why it works
  • Each tool has one clear job.
  • No full timeline edit is required just to test the idea.
  • You can get a usable result fast, then polish later.
Albert Einstein style talking-head thumbnail with a large expressive face and bold text about creating an AI avatar that can talk and move
1 Upgrade the still first

Start from a flat source image, then create a cleaner 3D-style version so the final talking clip looks more deliberate.

2 Create the voice in ElevenLabs

Export one clean MP3 first so the avatar step has a finished voice track to work from.

3 Build the talking avatar in HeyGen

Bring the image and audio together in a talking-photo workspace so the still turns into a speaking clip.

Best use

Classroom explainers, narrated slides, character intros, poster videos, or simple AI presenter tests.

Tool stack

ChatGPT or Gemini for prep, ElevenLabs for audio, HeyGen for the talking avatar, and PowerPoint or CapCut for optional polish.

Quick warning

This fast version does not solve subtitles or precise lip sync on its own.

The workflow

A clean talking-photo stack you can build in one sitting

The useful part of the original note is not just the free-tool angle. It is the sequence. First you improve the character image. Then you generate the voice. Then you bring both pieces into HeyGen so the avatar speaks from that finished audio. That is easier to control than asking one app to do everything at once.

Related watch Einstein process HeyGen + Leonardo 3m 29s

Watch a closer process match before you build

This Ernesto Kenji tutorial is a better fit because it actually walks through an Albert Einstein talking-avatar build with HeyGen and Leonardo.ai. That makes it more useful here than a pure showcase clip.

01

Make a cleaner version of the character image

Start with ChatGPT or Gemini to help restyle the character visual. The idea is to take the original flat image and rebuild it as a sharper, more presentation-ready still before animation begins. That extra step matters because motion exaggerates every flaw in the source frame.

Albert Einstein style AI avatar portrait prepared as the clean starting still for a talking-head workflow
Search target for this step: a clean Einstein AI avatar or portrait you can use as the starting still.
  • Use the original image as your base reference.
  • Restyle it into a cleaner, more presentation-ready character.
  • Keep the expression and pose simple so the next step stays stable.
02

Generate the voice first so the avatar has something real to follow

Once the character still is ready, write the line and make the voice track in ElevenLabs. Export the MP3 before you animate anything. That gives you a fixed piece of audio to build around instead of guessing timing later.

Branded ElevenLabs text to speech page with the ElevenLabs name and Text to Speech heading visible
Branded tool reference for this step: the official ElevenLabs Text to Speech page, with the ElevenLabs name visible in the screenshot.
  • Paste the script into the TTS tool and export a clean MP3.
  • Keep the read simple and clear so the avatar step stays believable.
  • Lock the voice first, then animate to that audio instead of changing both at once.
03

Bring the image and audio into HeyGen to make the talking avatar

Now move the upgraded still and the finished MP3 into HeyGen. Its talking-photo workflow is the integration step: one image plus one voice track becomes a short speaking avatar clip. If you want, you can drop the exported result into PowerPoint afterward, but the actual talking-avatar build happens here.

Branded HeyGen talking photo page with the HeyGen name and the talking photo product heading visible
Branded tool reference for this step: the official HeyGen talking photo page, with the HeyGen name visible in the screenshot.
  • Upload the upgraded still as the avatar input.
  • Add the finished MP3 so the mouth movement follows your real voice track.
  • Export the clip, then use PowerPoint or CapCut only if you want presentation timing or extra polish.

What makes this workflow useful

It is fast, approachable, and good enough for real use. If you need a talking visual for a lesson, explainer, or lightweight AI presentation, this gets you there without opening a heavy video project.

Where it starts to break

The note is honest about the tradeoff: the simple version does not automatically give you subtitles or perfect mouth-to-word matching. If that part matters, you still need an extra polish pass.

Best ways to use it

You do not have to stop at textbook characters. The same structure works for any situation where one character or one image needs to deliver a short line clearly.

  • turning a mascot into a short presenter clip
  • making a historical or illustrated figure speak in an explainer
  • building narrated slides that feel less static
  • testing a talking-avatar concept before committing to a full edit

Common snag to fix early

Several comments under the note point to the same issue: short clips can feel too brief or slightly unsynced. A practical fix from the creator is to let the video loop inside PowerPoint. For captions or closer mouth sync, move the assembled clip into CapCut after the first pass.