bitHuman for developers

Give your app a face

bitHuman is a real-time avatar SDK: send it the speech your companion says and it returns a lip-synced face — a photoreal person with Essence 2 or any character with Expression 2 — rendered on iPhone, iPad, Android, Mac, Linux or in the browser.

You bring the persona, the voice and the language model; when the avatar renders in your app on the device and you use your own voice and language services, bitHuman receives usage metering only, never audio, video or conversation text.

Pick a platform below. Each has a copy-paste quickstart in the docs and a working example you can clone.

Deploying in a bank, a hospital or on a show floor? Enterprise →

Platforms

Pick your platform

bitHuman's SDKs add a real-time, lip-synced avatar to iPhone, iPad, Mac, Android, desktop (macOS and Linux) and web apps. The Swift and Android SDKs only render: your app passes in 16 kHz mono speech from any voice stack and draws the frames they return, so the persona, the voice and the language model are yours to choose. The web embed can also run the whole conversation with bitHuman's managed agent. From 12 October 2026, API and SDK use requires the Creator plan or higher. Rendering on the device uses 2 credits per minute of active session time, talking or idle.
  • iPhone & iPad

    A Swift package that renders the avatar on the iPhone or iPad.

    .package(url: "https://github.com/bithuman-product/homebrew-bithuman.git", from: "2.19.0")

    Quickstart Docs

  • Android

    Two Maven Central artifacts that render the avatar on arm64 phones.

    implementation("ai.bithuman:expression2-android:0.5.2")   // any character
    implementation("ai.bithuman:essence2-android:0.8.1")      // a photoreal person

    Quickstart Docs

  • Mac

    The Swift package for Mac apps, or the CLI for a desktop companion.

    .package(url: "https://github.com/bithuman-product/homebrew-bithuman.git", from: "2.19.0")

    Quickstart Docs

  • Web

    An iframe or a floating widget, in any site or React app.

    <iframe src="https://www.bithuman.ai/embed/A23WJF0199" allow="microphone *" style="width:100%;height:600px;border:0"></iframe>

    Quickstart Docs

  • Python

    Load an avatar, push audio and get frames on macOS or Linux.

    pip install "bithuman[expression-2]"

    Quickstart Docs

There is no native Windows SDK. On Windows, the Python SDK and the CLI run under WSL2, and the web embed runs in the browser.

Architecture

How an app fits together

The Swift and Android SDKs only render: your app passes in 16 kHz mono speech and draws the frames they return. The conversation — speech recognition, the language model and the voice — comes from services you choose, or from bitHuman's managed agent through the web embed. On a Mac or a Linux PC, the bitHuman CLI's local conversation brain can run all three on the machine; the session still reports usage online.

The avatar renders

On the device: iPhone, iPad, Mac or Android
Or in the bitHuman cloud (the web embed, LiveKit)

16 kHz mono speech in, lip-synced frames out

The conversation runs

In your own speech, language-model and voice services
Or in bitHuman's managed agent (the web embed)

Your own voice stack

The avatar renders
On the device (iPhone, iPad, Mac or Android), or on your Mac or Linux machine with Python
The conversation runs
Your speech recognition, language model and voice
What reaches bitHuman
Usage metering only — never audio, video or conversation text
Use
Swift package, Android SDK, Python

bitHuman's managed agent

The avatar renders
In the bitHuman cloud, or in the browser tab with WebGPU
The conversation runs
bitHuman's servers, with your persona and, if you like, your own model or voice keys
What reaches bitHuman
The conversation's audio and text, to run it
Use
Web embed and website widget

The local conversation brain

The avatar renders
On a Mac or a Linux PC
The conversation runs
On the same machine (BITHUMAN_LOCAL=1)
What reaches bitHuman
Usage metering only; the session still reports usage online
Use
The bitHuman CLI

Build a companion app in the docs

Performance

Measured on real devices

Every configuration we publish here renders faster than real time, including 10-minute runs on the phones.
Seconds of avatar video rendered per second, on phones, a Mac and in the browser
DeviceEssence 2Expression 2
macOS · CLIApple M44.24× real time2026-09-278.4× real time2026-09-27
macOS · Swift packageApple M44.8× real time2026-09-248.85× real time2026-09-24
iPhone · Swift packageiPhone 152.16× real time2026-09-275.55× real time2026-09-27
AndroidSamsung Galaxy S25+2.08× real time2026-09-252.4× real time2026-09-23
Web browser (WebGPU)Chrome on Apple M41.72× real time2026-09-271.95× real time2026-09-27
iPhone · Swift package · held 10 miniPhone 151.32× real time2026-09-255.15× real time2026-09-25
Android · held 10 minSamsung Galaxy S25+1.48× real time2026-09-272.2× real time2026-09-27
Web browser (WebGPU) · held 10 minChrome on Apple M42.16× real time2026-09-272.05× real time2026-09-25
“× real time” is seconds of avatar video rendered per second of wall-clock time: end-to-end audio in to frame out, one session, unpaced; the slowest of three quiet runs at least ten minutes apart on the published release. Read from docs performance.json (generated 2026-09-27). How we measure →

Pricing

What it costs to run an avatar in your app

What it costs to run an avatar in your app
ItemOn the device (Essence 2, Expression 2)bitHuman cloud avatarManaged voice chat (all-inclusive)
Credits per minute of active session time2410
At top-up rates ($1 = 100 credits)about $0.02about $0.04about $0.10
Minutes in a Creator month (1,800 credits)900450180
Speech, language model and voiceYou pay your providersYou pay your providersIncluded
Concurrent sessionsLimited by creditsPer plan (Creator 3 … Enterprise 200)Per plan (Creator 3 … Enterprise 200)
  • From 12 October 2026, API and SDK use requires the Creator plan or higher. On a paid plan you can top up at any time.
  • Idle is billed. A companion that stays on screen bills for its whole active session, talking or idle. End the session when the user leaves, and show a still frame instead of an idle loop when nobody is talking.
  • Worked example. A user whose on-device sessions add up to 20 minutes a day uses about 1,200 credits in 30 days.
  • Your own avatar. Creating one costs 500 credits for Essence 2 or 2,000 credits for Expression 2. A Creator month is 1,800 credits.

Rates from GET https://api.bithuman.ai/v1/pricing, as published on docs.bithuman.ai/pricing.

Privacy

What reaches bitHuman

When the avatar renders in your app on the device and you use your own voice and language services, bitHuman receives usage metering only, never audio, video or conversation text. With the web embed, the conversation runs on bitHuman's servers, even when the avatar renders in the tab.

Security & privacy

Checklist

Before you ship

  1. Plan

    From 12 October 2026, API and SDK use requires the Creator plan or higher.

    Plans and top-ups
  2. Credential

    The Apple and Android SDKs authenticate with your account's API secret today, so a shipped app carries it. Fetch it from your backend at startup, never compile it in, use a separate secret for each app, and rotate it if it leaks.

    What a shipped app holds (docs)
  3. AI disclosure

    Tell your users they are talking to an AI.

    Acceptable use

Examples

Examples you can clone

  • iOS — Expression 2: a still from the recording

    iOS — Expression 2

    A SwiftUI app with a talking character, rendered on the iPhone or iPad.

    Runs on
    iPhone, iPad
    Conversation
    No: it speaks a bundled clip, or lip-syncs your microphone
    Needs
    API secret, Xcode 26, an iPhone or iPad

    swift/ios-expression2 Walkthrough

  • Mac — Expression 2: a still from the recording

    Mac — Expression 2

    One Swift command-line tool: speech in, lip-synced frames out, on your Mac.

    Runs on
    Mac (Apple silicon)
    Conversation
    No: it renders a speech file
    Needs
    API secret, Xcode 26

    swift/macos-expression2 Walkthrough

  • Android — Expression 2 and Essence 2: a still from the recording

    Android — Expression 2 and Essence 2

    Two complete Gradle apps that render a talking avatar on the phone.

    Runs on
    Android arm64
    Conversation
    No: each renders a speech clip
    Needs
    API secret, a physical arm64 phone

    android/expression2-hello android/essence2-hello Walkthrough

  • Python quickstart: a still from the recording

    Python quickstart

    Open an avatar, play speech through it and watch it talk; conversation.py adds a voice conversation.

    Runs on
    Mac, Linux
    Conversation
    Yes, in conversation.py (with your own OpenAI key)
    Needs
    API secret

    python/quickstart Walkthrough

  • Screenshot coming

    Next.js

    A browser front end for an avatar agent in a LiveKit room.

    Runs on
    Web
    Conversation
    Yes, through the agent in the room
    Needs
    LiveKit and an avatar agent (which needs an API secret)

    integrations/nextjs-ui

  • CLI recipes: a still from the recording

    CLI recipes

    Render a video or run a live avatar from the terminal, with no code.

    Runs on
    Mac, Linux
    Conversation
    Yes, with bithuman run
    Needs
    bithuman login

    api/cli Walkthrough

All examples

AI companions

Building a companion?

The companion guide walks through a companion screen in an iPhone, iPad or Android app: the avatar renders on the phone and speaks the replies your own voice stack produces — the SDK, the API secret, resampling speech, replies, interruption, and closing the avatar with the screen.

Read the guide

For AI coding agents

Docs your coding agent can read

FAQ

Questions developers ask

Is there an iOS and Android SDK for bitHuman avatars?

Yes. The Swift package renders Essence 2 and Expression 2 on iPhone, iPad and Mac. essence2-android and expression2-android on Maven Central render them on arm64 Android phones. Measured results for each device are published at docs.bithuman.ai/performance.

Does the avatar render on the phone or in the cloud?

With the Swift and Android SDKs, on the phone. With the web embed, in the bitHuman cloud by default, or in the browser tab with WebGPU (render=local), falling back to cloud rendering. With a LiveKit cloud avatar, in the bitHuman cloud.

Does the SDK include the voice and the AI conversation?

The Swift and Android SDKs only render: your app passes in 16 kHz speech and draws the frames. Use any speech recognition, language model and voice, or use bitHuman's managed agent through the web embed. On a Mac or a Linux PC, the bitHuman CLI's local conversation brain can run the whole conversation on the machine.

Can I use my own language model and persona?

Yes. Any OpenAI-compatible endpoint works. Set the persona in your model's system prompt or, for a managed agent, in its system_prompt.

Can I build an AI companion app with bitHuman?

Yes. bitHuman provides the companion's face, rendered in real time on the phone, on the Mac or in the browser, and you choose the voice, the language model and the persona. The companion guide in the docs walks through adding the SDK, the API secret, resampling speech, replies, interruption and closing the avatar with the screen.

What does it cost?

From 12 October 2026, API and SDK use requires the Creator plan or higher. Rendering on the device uses 2 credits per minute of active session time, talking or idle, which is about $0.02 at top-up rates. A bitHuman cloud avatar uses 4 credits per minute, and the all-inclusive managed voice chat uses 10.

Does it work offline on a phone?

No. Apps on phones and Macs need a connection to start a session, and they keep rendering through a network drop of up to 5 minutes. The web embed needs a connection throughout. Offline license is only available to Business and Enterprise clients who want to run realtime avatars completely locally, off the internet — e.g. kiosks, trade shows, ATM machines, embedded screens. Linux PCs and terminals; arranged through sales.

What data does the SDK send to bitHuman?

Usage reports for metering: which session ran and for how long. They never contain audio, video, images or conversation text. The SDK also checks your API secret when a session starts. With the web embed, the conversation itself runs on bitHuman's servers.

Can I use bitHuman in a React or Next.js app?

Yes, with the iframe embed or the floating website widget. There is no npm package: the embed is an iframe in any framework, and the docs web page has a React snippet.

All questions