Ex-Omni-2D — omni-modal dialogue → talking avatar

Ex-Omni-2D answers you in one shot: an 11B omni-modal LLM writes a Visual Thought Plan (<avatar_plan>: scene, emotion, movement style, motion description) plus a spoken reply, synthesises that reply in the reference speaker's voice with Qwen3-TTS, and feeds the resulting speech tokens to an OmniAvatar / Wan2.1-T2V-1.3B diffusion generator that renders the avatar.

Give it a portrait, a short voice sample to clone, and something to say.

Examples (reference portrait and voice from the Ex-Omni-2D repo)