Add a talking avatar to your LiveKit agent
What the plugin does
livekit-plugins-bithuman turns your agent's reply audio into lip-synced video of an avatar. The avatar joins the room as a participant, and your agent keeps its own speech recognition, language model and voice.
It works on LiveKit Cloud or your own LiveKit server. It is a Python plugin; there is no Node.js plugin. You choose where the avatar renders when you start the session:
- Pass
avatar_id=with an agent code, and the avatar renders in the bitHuman cloud, in the US. bitHuman publishes its audio and video into your room. - Pass
model_path=with an avatar file, and the avatar renders inside your worker's process. Its audio and video stay on your machine until your worker publishes them.
Install
You need Python 3.10 to 3.14, a LiveKit project with its URL and credentials, a bitHuman API secret, and an OpenAI key for the voice model in the example.
To render the avatar inside your worker with model_path=, also install "bithuman[expression-2]".
pip install "livekit-agents[openai,silero]" livekit-plugins-bithuman python-dotenvKeep your API secret out of the room
In the worker, name your API secret BITHUMAN_MASTER_SECRET. For a cloud avatar, the worker mints a one-hour token that can start only this agent's avatar in this room, and passes it to the plugin as api_secret=.
Keep BITHUMAN_API_SECRET unset in the worker. Without api_secret=, the plugin reads that variable and copies it into the avatar's participant attributes, which everyone in the room can read.
Start the avatar beside your agent
These are the avatar lines of the worker. The complete agent.py, with the token mint call, is on the docs page.
Run python agent.py dev, then join the room from the LiveKit Agents Playground. The avatar appears and answers with its lips in sync. If two voices play, the agent session is also publishing audio: keep audio_output=False so only the avatar speaks.
avatar = bithuman.AvatarSession(
avatar_id=agent_code,
api_secret=await livekit_cloud_token(agent_code, ctx.room.name),
)
await avatar.start(session, room=ctx.room)
await session.start(
agent=Agent(instructions="You are a friendly assistant. Keep answers short."),
room=ctx.room,
room_options=RoomOptions(audio_output=False), # the avatar publishes the audio
)Where the audio goes
The avatar does not listen or think: it renders the speech your agent already produces. With a cloud avatar, the audio reaches bitHuman and the avatar renders in the bitHuman cloud, in the US. When it renders in your worker, its audio and video stay with you, and bitHuman receives a credential check and usage reports.
Choose a model
Either way, the plugin serves the agent's own model, Essence 2 or Expression 2. Essence 2 renders a photoreal person from one portrait. Expression 2 renders any character, from people to animals and cartoons, from one portrait.
While you build, use the wise-pup sample (Expression 2, agent code A23WJF0199) or sofia-ramirez (Essence 2, A52DHS2219), or create your own avatar from one portrait.
What it costs
From 12 October 2026, API and SDK use requires the Creator plan or higher.
A cloud avatar bills the cloud rate, and an avatar in your worker bills the self-hosted rate. A session bills active session time, talking or idle, to the second.
Every step, with troubleshooting, is in the docs. LiveKit guide in the docs
