Developers · Android

A real-time avatar for your Android app

Two artifacts on Maven Central — essence2-android for a photoreal person and expression2-android for any character — render the avatar on the handset. After the one-time model download, the SDK's only network traffic is usage reporting.

From 12 October 2026, API and SDK use requires the Creator plan or higher.

Install

Install

Gradle (Maven Central)
implementation("ai.bithuman:expression2-android:0.5.2")   // any character
implementation("ai.bithuman:essence2-android:0.8.1")      // a photoreal person

What you need: arm64 device, API secret (from the docs)

Your first frame, step by step The full Android guide

Conversation

Add a conversation

The Android SDK only renders: your app passes in 16 kHz mono speech and draws the frames it returns. Bring your own voice stack, or let bitHuman's managed agent run the conversation instead.
  • Your own speech recognition, language model and voice

    Any stack works: the SDK takes 16 kHz speech. OpenAI Realtime returns 24 kHz audio: resample it to 16 kHz before you pass it in. The converter is in the docs.

    Resample speech to 16 kHz

  • Build the companion screen

    The companion guide walks through the whole loop on the phone, with Kotlin for each step: the SDK, the API secret, resampling, replies, interruption and closing the avatar with the screen.

    Companion app (docs)

  • Or let bitHuman run the conversation

    The web embed runs a managed agent on bitHuman's servers with your persona and, if you like, your own OpenAI-compatible model; a LiveKit agent can use a bitHuman cloud avatar. Either way the avatar renders in the bitHuman cloud (the web embed's default), not on the phone.

    Web embed

When the avatar renders in your app on the device and you use your own voice and language services, bitHuman receives usage metering only, never audio, video or conversation text.

Requirements

Requirements

  • A physical arm64 phone; emulators cannot load the engines.
  • The two artifacts take different audio formats and minimum SDK levels; the docs compare them.

Performance

Measured on real hardware

Every configuration we publish for this platform renders faster than real time.
Seconds of avatar video rendered per second, Android
DeviceEssence 2Expression 2
AndroidSamsung Galaxy S25+2.08× real time2026-09-252.4× real time2026-09-23
Android · held 10 minSamsung Galaxy S25+1.48× real time2026-09-272.2× real time2026-09-27
“× real time” is seconds of avatar video rendered per second of wall-clock time: end-to-end audio in to frame out, one session, unpaced; the slowest of three quiet runs at least ten minutes apart on the published release. Read from docs performance.json (generated 2026-09-27). How we measure →

Pricing

What it costs

What it costs to run an avatar in your app
ItemOn the device (Essence 2, Expression 2)bitHuman cloud avatarManaged voice chat (all-inclusive)
Credits per minute of active session time2410
At top-up rates ($1 = 100 credits)about $0.02about $0.04about $0.10

From 12 October 2026, API and SDK use requires the Creator plan or higher.

Rates from GET https://api.bithuman.ai/v1/pricing, as published on docs.bithuman.ai/pricing.

The full table, with a worked example Budget an app (docs)

Examples

Examples you can clone

All examples

Checklist

Before you ship

The full checklist

Limits

Not available on this platform

  • A conversation brain that runs on the phone: the conversation is your app's, or runs in the cloud.
  • Offline operation: a session needs a connection to start, and keeps rendering through a network drop of up to 5 minutes.

More

Other platforms