The avatar renders on the device or on your own servers, so its audio and video stay there. Speech recognition, the language model and the voice run where you point them: the CLI's local conversation brain on a Mac or Linux machine, or your own services inside your network. bitHuman then receives a credential check when a session starts, the avatar download, and usage reports with no audio, video or conversation text.
Creating an avatar from a portrait happens in the bitHuman cloud; the finished avatar model then runs on your hardware.