Ex-Omni-2D — omni-modal dialogue → talking avatar
Ex-Omni-2D answers you in
one shot: an 11B omni-modal LLM writes a Visual Thought Plan
(<avatar_plan>: scene, emotion, movement style, motion description) plus a
spoken reply, synthesises that reply in the reference speaker's voice with
Qwen3-TTS, and feeds the resulting speech tokens to an OmniAvatar /
Wan2.1-T2V-1.3B diffusion generator that renders the avatar.
Give it a portrait, a short voice sample to clone, and something to say.
Examples (reference portrait and voice from the Ex-Omni-2D repo)