Deployment options

Four ways to run a bitHuman avatar

Every mode uses the same avatar. What changes is where the avatar renders, where the conversation runs, and what reaches us.

The modes, side by side

bitHuman cloud

The avatar renders
On bitHuman's servers, in the US
The conversation runs
bitHuman's voice service, or your own provider keys
What reaches bitHuman
The session's audio and conversation, to run it. Transcripts are kept with your agent; deleting the agent deletes them.
Internet
Required for the whole session
Models
Essence 2, Expression 2 and the first-generation models
Best for
The fastest start: web pages, apps and APIs

bitHuman cloud in the docs

Your servers

The avatar renders
On your own Mac or Linux machines, on-premises or in your cloud account; a standard Linux PC needs no GPU
The conversation runs
Your choice: the CLI's local conversation brain, your own speech and language services, or bitHuman's
What reaches bitHuman
A credential check when a session starts, the avatar download, and usage reports with no audio, video or conversation text
Internet
To start a session; rendering continues through a network drop of up to 5 minutes
Models
Essence 2 and Expression 2; Essence 1 on the CLI and Python
Best for
On-premises, data-center and private-cloud deployments

Your servers in the docs

On the device

The avatar renders
On the iPhone, iPad, Mac or Android device in front of the user, or in a WebGPU browser tab
The conversation runs
Your app's choice; with the web embed, on bitHuman's servers
What reaches bitHuman
With your own voice and language services, usage metering only — never audio, video or conversation text
Internet
To start a session; rendering continues through a network drop of up to 5 minutes
Models
Essence 2 and Expression 2
Best for
Apps, kiosks and embedded screens

On the device in the docs

Fully offline

The avatar renders
On your Linux PCs and terminals
The conversation runs
Agreed with sales for your site
What reaches bitHuman
Nothing while it runs: usage is metered on the machine, and no reconnection is required
Internet
Off the internet; creating the avatar happens online first
Models
Essence 1 now · Essence 2 & Expression 2 later
Best for
Kiosks, ATM machines, trade shows and embedded screens — Business and Enterprise

Fully offline in the docs

Creating an avatar from a portrait always happens in the bitHuman cloud. Once created, its model file downloads to your machines and runs there.

Fully offline

Offline license is only available to Business and Enterprise clients who want to run realtime avatars completely locally, off the internet — e.g. kiosks, trade shows, ATM machines, embedded screens. Linux PCs and terminals; arranged through sales.

Essence 1 runs fully offline on Linux (x86_64 and ARM64) today. Essence 2 and Expression 2 offline come later. Expression 1 runs in the bitHuman cloud only.

Offline licensing

Which should I choose?

  • Want the conversation to stay on your network?

    Choose Your servers or On the device, with the CLI's local conversation brain or your own speech and language services.

    See the mode

  • No reliable internet at the site?

    Choose Fully offline (Business and Enterprise).

    See the mode

  • Just building a web experience?

    Start with the bitHuman cloud.

    See the mode

Go deeper