Self hosted, real-time digital human agent platform. Build voice-first AI agents with WebRTC, persona memory, tools, RAG, and optional digital-human video.
-
Updated
Aug 5, 2026 - Python
Self hosted, real-time digital human agent platform. Build voice-first AI agents with WebRTC, persona memory, tools, RAG, and optional digital-human video.
Talking Head (3D): A JavaScript class for real-time lip-sync using full-body 3D avatars.
Local-first browser AI video editor with WebGPU AI music, AI repair, voiceovers, captions, talking avatars, and deterministic export.
🎭 AI Avatar / digital human platform — upload a photo, clone a voice, talk to any face in real time with lip-sync video. Open-source, self-hosted. Claude · Whisper · Chatterbox · MuseTalk.
[CVPR-2025] The official code of HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation
DialogLab is an authoring tool for configuring and running Human-AI multi-party conversations.
[NeurlPS-2024] The official code of MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
Explore the power of Azure Text-to-Speech with interactive talking avatar, Lisa 👩🏻🦱. Choose from multiple languages and avatar styles to bring your text to life.
ComfyUI custom nodes for LongCat Video Avatar 1.5 audio-driven human video generation; a macOS inference branch, macOS-MLX, is now available for Apple Silicon MLX testing.
Interactable AI that have control over your frontend website, It guides your user walk around your website its a salesman / supports.
Talking Avatar: create video from plain text or audio file in minutes, support up to 100+ languages and 350+ voice models.
Budget-aware AI agent content studio for creating videos, images, carousels, voiceovers, music, captions, and content calendars with consistent brand assets and cost control.
AI Avatar/Anchor: create video from plain text or audio file in minutes, support up to 100+ languages and 350+ voice models.
The text to speech avatar system is a text to speech feature with vision capabilities, that allow customers to create synthetic videos of a 2D photorealistic avatar speaking. The Neural text to speech Avatar models are trained by deep neural networks based on the human video recording samples, and the voice of the avatar .
Animated Characters: create video from plain text or audio file in minutes, support up to 100+ languages and 350+ voice models.
The official main page of "EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion".
Audio-driven talking avatar and lip-sync video generation from a reference image and speech audio.
These are output demos of different models for Talking Avatar Generation (TAG).
Digital-human / talking-avatar workspace orchestrating InfiniteTalk, MuseTalk, and Qwen3-TTS for audio-driven portrait video generation.
Talking avatars created using Leonardo.ai, VoiceOverMaker, and D-ID.
Add a description, image, and links to the talking-avatar topic page so that developers can more easily learn about it.
To associate your repository with the talking-avatar topic, visit your repo's landing page and select "manage topics."