Real-time talking avatars in your Python code
What the SDK does
The bithuman package renders a talking avatar from audio. On Linux and Windows it renders both models on the CPU alone, with no GPU. On macOS it renders on Apple silicon. Frames come back as RGB numpy arrays, so you can show them, encode them or send them anywhere.
Install
You need:
- Python 3.10 to 3.14.
- macOS 14 or later on Apple silicon, Linux on x86_64 or arm64, or Windows 11 on x86_64.
- A bitHuman API secret.
- About 1 GB of disk for the package and an avatar.
python3 -m venv .venv
source .venv/bin/activate
pip install "bithuman[expression-2]"Render your first avatar
Set BITHUMAN_API_SECRET, download the wise-pup sample and a sample speech file, and render. out_mp4= needs no ffmpeg. render takes a path to an audio file, such as WAV or MP3, or decoded 16 kHz mono audio.
export BITHUMAN_API_SECRET="<your API secret>"
curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2"
curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wavimport bithuman
bithuman.open("wise-pup.imx").render("speech.wav", out_mp4="out.mp4")Work with the frames
To process the frames yourself, iterate over render without out_mp4=. Frames are RGB; OpenCV expects BGR, so reverse the channels before you write one with OpenCV.
import bithuman
with bithuman.open("wise-pup.imx") as avatar:
frames = [image for image in avatar.render("speech.wav")]
print(len(frames), "frames of", frames[0].shape)Stream live audio
For a live conversation, AsyncBithuman takes each chunk of reply audio with its sample rate and yields frames with the matching audio, paced at the model's play rate. Call flush() when a reply ends and interrupt() when the user talks over it. The same code takes audio from OpenAI Realtime or any TTS service.
Common fixes
The avatar renders on your machine, so its audio and video stay with you; bitHuman receives a credential check and usage reports. If the first render fails:
externally-managed-environment: create and activate a virtual environment.NotSupportedopening an Expression 2 file: install"bithuman[expression-2]".NotAuthorised: setBITHUMAN_API_SECRETin this shell; nothing is rendered or written without it.- Frames look blue: your display wants BGR, so reverse the channels.
- Raw audio plays slow and long: pass a file path, or resample to 16 kHz mono.
Choose a model
One install renders both models, and the same call opens Essence 2 and Essence 1 avatar files. Essence 2 renders a photoreal person from one portrait. Expression 2 renders any character, from people to animals and cartoons, from one portrait.
While you build, use the wise-pup sample (Expression 2, agent code A23WJF0199) or sofia-ramirez (Essence 2, A52DHS2219), or create your own avatar from one portrait.
What it costs
From 12 October 2026, API and SDK use requires the Creator plan or higher.
A live session bills per second while the avatar runs, talking or idle. Rendering a file bills the length of the video it writes, at the self-hosted rate.
Every step, with troubleshooting, is in the docs. Python guide in the docs
