Skip to content

Give your Pipecat voice bot a talking avatar

Put one video service after your TTS: your bot's speech goes in, and lip-synced avatar video with the matching audio comes out.

What the integration does

pipecat-bithuman is a community integration for Pipecat, maintained by bitHuman, not by the Pipecat team. Its BitHumanVideoService takes the bot's speech and returns avatar video with the matching audio, paired frame by frame.

The avatar renders inside your bot's own process through the bitHuman Python SDK, on your own machine. Your pipeline keeps its own speech-to-text, LLM and TTS services. The avatar opens on StartFrame and closes on EndFrame, CancelFrame or cleanup.

Install

You need:

  • Python 3.11 to 3.14.
  • macOS 14 or later on Apple silicon, or Linux on x86_64 or arm64. A PC with no GPU renders both models.
  • A bitHuman API secret.
Install
python3 -m venv .venv
source .venv/bin/activate
pip install "pipecat-bithuman[expression-2]"

Add the avatar to your pipeline

Set BITHUMAN_API_SECRET and BITHUMAN_MODEL_PATH, turn on video_out_enabled=True in your transport, and add the service between tts and transport.output(). The service never logs the secret and removes it from error text.

If the bot speaks with no video and no ErrorFrame, video out is off in the transport.

bot.py (excerpt)
avatar = BitHumanVideoService()  # BITHUMAN_MODEL_PATH + BITHUMAN_API_SECRET
# …
pipeline = Pipeline(
    [
        transport.input(),
        stt,
        aggregators.user(),
        llm,
        tts,
        avatar,
        transport.output(),
        aggregators.assistant(),
    ]
)

The complete bot

Try it without a bot

The package's demo script sends a WAV file through a Pipecat pipeline, as a TTS service would, and writes the avatar frames that come out to an MP4. Add --interrupt-at 10 --repeat 2 to see a reply cut off and a new one follow.

Render the demo
export BITHUMAN_API_SECRET="<your API secret>"
git clone https://gitlab.com/bithuman/sdk/pipecat-bithuman
cd pipecat-bithuman
curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2"
curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav
BITHUMAN_MODEL_PATH=wise-pup.imx python examples/render_demo.py speech.wav demo.mp4

Where the audio goes

The avatar renders in your bot, so its audio and video stay with you. bitHuman receives a credential check and usage reports.

Common fixes

Most first-run problems show up as an ErrorFrame that names the cause:

  • No avatar model: pass model_path= or set BITHUMAN_MODEL_PATH.
  • The bitHuman Python SDK is not installed: activate the virtual environment and install "pipecat-bithuman[expression-2]" again.
  • The API secret is missing or was not accepted: set BITHUMAN_API_SECRET, or create a new secret.
  • The bot speaks but shows no video after an ErrorFrame: the avatar failed and the TTS audio passed through. Fix what the frame names.
  • Session time keeps running after the user leaves: cancel the pipeline, or end it with EndFrame.

Choose a model

The one install renders both models from an avatar file (.imx). Essence 2 renders a photoreal person from one portrait. Expression 2 renders any character, from people to animals and cartoons, from one portrait.

While you build, use the wise-pup sample (Expression 2, agent code A23WJF0199) or sofia-ramirez (Essence 2, A52DHS2219), or create your own avatar from one portrait.

What it costs

From 12 October 2026, API and SDK use requires the Creator plan or higher.

Usage bills per second while the avatar runs, talking or idle. The rate for your model is on the pricing page.

See pricing

Give your agent a face