Skip to content

Real-time talking avatars in your Python code

A file in and frames out, or a live stream of audio in and frames out, on your own machine.

What the SDK does

The bithuman package renders a talking avatar from audio. On Linux and Windows it renders both models on the CPU alone, with no GPU. On macOS it renders on Apple silicon. Frames come back as RGB numpy arrays, so you can show them, encode them or send them anywhere.

Install

You need:

  • Python 3.10 to 3.14.
  • macOS 14 or later on Apple silicon, Linux on x86_64 or arm64, or Windows 11 on x86_64.
  • A bitHuman API secret.
  • About 1 GB of disk for the package and an avatar.
Install
python3 -m venv .venv
source .venv/bin/activate
pip install "bithuman[expression-2]"

Render your first avatar

Set BITHUMAN_API_SECRET, download the wise-pup sample and a sample speech file, and render. out_mp4= needs no ffmpeg. render takes a path to an audio file, such as WAV or MP3, or decoded 16 kHz mono audio.

Shell
export BITHUMAN_API_SECRET="<your API secret>"
curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2"
curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav
Python
import bithuman
bithuman.open("wise-pup.imx").render("speech.wav", out_mp4="out.mp4")

Work with the frames

To process the frames yourself, iterate over render without out_mp4=. Frames are RGB; OpenCV expects BGR, so reverse the channels before you write one with OpenCV.

Python
import bithuman

with bithuman.open("wise-pup.imx") as avatar:
    frames = [image for image in avatar.render("speech.wav")]
print(len(frames), "frames of", frames[0].shape)

Stream live audio

For a live conversation, AsyncBithuman takes each chunk of reply audio with its sample rate and yields frames with the matching audio, paced at the model's play rate. Call flush() when a reply ends and interrupt() when the user talks over it. The same code takes audio from OpenAI Realtime or any TTS service.

Common fixes

The avatar renders on your machine, so its audio and video stay with you; bitHuman receives a credential check and usage reports. If the first render fails:

  • externally-managed-environment: create and activate a virtual environment.
  • NotSupported opening an Expression 2 file: install "bithuman[expression-2]".
  • NotAuthorised: set BITHUMAN_API_SECRET in this shell; nothing is rendered or written without it.
  • Frames look blue: your display wants BGR, so reverse the channels.
  • Raw audio plays slow and long: pass a file path, or resample to 16 kHz mono.

Choose a model

One install renders both models, and the same call opens Essence 2 and Essence 1 avatar files. Essence 2 renders a photoreal person from one portrait. Expression 2 renders any character, from people to animals and cartoons, from one portrait.

While you build, use the wise-pup sample (Expression 2, agent code A23WJF0199) or sofia-ramirez (Essence 2, A52DHS2219), or create your own avatar from one portrait.

What it costs

From 12 October 2026, API and SDK use requires the Creator plan or higher.

A live session bills per second while the avatar runs, talking or idle. Rendering a file bills the length of the video it writes, at the self-hosted rate.

See pricing

Give your agent a face