Add a talking avatar to OpenAI Realtime
How it fits together
The avatar does not listen or think: it renders the speech your agent already produces. OpenAI Realtime keeps the conversation, and bitHuman turns each reply into video of an avatar, with the audio in step.
In Python, AsyncBithuman takes each chunk of reply audio with its sample rate and yields frames with the matching audio. It renders on your Mac, Linux or Windows machine. In an iPhone, iPad, Mac or Android app, the Swift package or the Android SDK takes the same audio, resampled to 16 kHz mono first.
Install
Install the Python SDK into a virtual environment, then download a sample avatar file.
python3 -m venv .venv
source .venv/bin/activate
pip install "bithuman[expression-2]"
curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2"Push the reply audio
Feed AsyncBithuman the PCM audio your stack plays, from OpenAI Realtime or a TTS service. Call flush() when a reply ends and interrupt() when the user talks over it. Until a reply arrives you get idle frames; then the lips follow its words.
import asyncio, soundfile as sf
from bithuman import AsyncBithuman
async def main():
avatar = await AsyncBithuman.create(model_path="wise-pup.imx") # reads BITHUMAN_API_SECRET
pcm, rate = sf.read("speech.wav", dtype="int16")
async def speak():
for i in range(0, len(pcm), rate // 10): # 100 ms chunks, as they arrive
await avatar.push_audio(pcm[i:i + rate // 10].tobytes(), rate, last_chunk=False)
await avatar.flush() # end of the reply
task = asyncio.create_task(speak())
try:
async for frame in avatar.run(): # paced at the model's play rate
if frame.has_image:
show(frame.bgr_image) # BGR numpy array
if frame.audio_chunk:
play(frame.audio_chunk.array) # audio in sync with the frame
finally:
task.cancel()
await avatar.shutdown() # frees the model and the credential
asyncio.run(main())Run the whole conversation
The quickstart in the examples repository runs the full loop in a desktop window: your microphone goes to OpenAI Realtime, the reply's audio goes into the avatar with push_audio and flush, and lip-synced frames and audio come out. It needs your API secret and an OpenAI key.
In a LiveKit Agents worker, the bitHuman plugin pairs the avatar with openai.realtime.RealtimeModel directly.
Where the audio goes
The reply audio goes from your code into the avatar on your machine. The avatar's audio and video stay with you; bitHuman receives a credential check and usage reports. The conversation itself is between your code and OpenAI, or the relay below.
No OpenAI key of your own
The bitHuman Realtime relay opens an OpenAI Realtime voice session through bitHuman and bills it in bitHuman credits. It speaks the OpenAI Realtime WebSocket protocol, unchanged, and authenticates with your API secret. Pair it with an on-device avatar to give the voice a face, as the Flutter plugin and the CLI do.
Choose a model
The Python SDK renders both second-generation models. Essence 2 renders a photoreal person from one portrait. Expression 2 renders any character, from people to animals and cartoons, from one portrait.
While you build, use the wise-pup sample (Expression 2, agent code A23WJF0199) or sofia-ramirez (Essence 2, A52DHS2219), or create your own avatar from one portrait.
What it costs
From 12 October 2026, API and SDK use requires the Creator plan or higher.
A session bills active session time, talking or idle, to the second. The relay bills its voice session in credits too.
Every step, with troubleshooting, is in the docs. Voice agent guide in the docs
