--- license: other license_name: ltx-2-community-license-agreement license_link: https://github.com/Lightricks/LTX-2/blob/main/LICENSE pipeline_tag: image-to-video base_model: - Lightricks/LTX-2.3 tags: - image-to-video - video-generation - audio-video-generation - reference-to-video - long-video - multi-shot - dmd ---

JoyAI-Echo generated video gallery

JoyAI-Echo 1.5

🎬 Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

📄 Paper | 🌐 Project Page | 💻 Inference Code | 🎬 Director Agent

Reference-to-Video Audio + Video 8-step DMD BF16, FP8 and FP4

## Model Summary **JoyAI-Echo 1.5** is a unified audio-visual generation system for persistent stories and interactive worlds. This repository contains the long-video model weights. Its reference-to-video pipeline accepts a text prompt, an optional first-frame condition, and up to seven ordered memory slots containing reference images and audio. A paired cross-modal memory bank carries character appearance, voice identity, and scene context across independently editable shots. The release uses an 8-step DMD generator built on LTX-2.3 and provides BF16, FP8, and FP4 variants through one inference pipeline. ## Highlights - 🎞️ **Reference-driven multi-shot generation:** use text, a first frame, and up to seven memory slots for each shot. - 🔊 **Joint audio-video generation:** generate synchronized video and audio in one pipeline. - 🧠 **Paired cross-modal memory:** reuse appearance, voice, and scene evidence from previous shots. - ⚡ **Few-step inference:** generate each shot with an 8-step DMD sampler. - 💻 **Consumer-GPU profiles:** layer-wise DiT offload and tiled Video VAE decoding are available in the inference repository. ## Released Checkpoints | Directory | Precision | Size | Contents | | --- | --- | ---: | --- | | `echo15_full_dmd/` | BF16 | 46.14 GB | `model.safetensors` | | `echo15_fp8/` | FP8 | 27.62 GB | `model.safetensors` | | `echo15_fp4/` | FP4 | 22.81 GB | `components.safetensors`, `transformer_modelopt.pt` | Each directory also contains a `checkpoint.json` manifest. `SHA256SUMS` records the hashes of all checkpoint files. JoyAI-Echo 1.0 is preserved on the [`echo1.0`](https://huggingface.co/jdopensource/JoyAI-Echo/tree/echo1.0) branch. ## Usage ### 1. Download one checkpoint ```bash # BF16 hf download jdopensource/JoyAI-Echo \ --include "echo15_full_dmd/*" \ --local-dir checkpoints # FP8 hf download jdopensource/JoyAI-Echo \ --include "echo15_fp8/*" \ --local-dir checkpoints # FP4 hf download jdopensource/JoyAI-Echo \ --include "echo15_fp4/*" \ --local-dir checkpoints ``` Download the Gemma text encoder separately: ```bash hf download google/gemma-3-12b-it \ --local-dir checkpoints/gemma-3-12b ``` ### 2. Get the inference code ```bash git clone https://github.com/jd-opensource/JoyAI-Echo.git cd JoyAI-Echo/echo_longvideo ``` ### 3. Run batch inference ```bash python inference.py --config configs/inference.bf16.yaml # BF16 python inference.py --config configs/inference.fp8.yaml # FP8 python inference.py --config configs/inference.fp4.yaml # FP4 ``` The default configuration processes the R2V requests in `examples/the_last_visa/requests/`. See the [inference repository](https://github.com/jd-opensource/JoyAI-Echo/tree/main/echo_longvideo) for the request schema, conditioning cache, low-VRAM profiles, and local server. ## Links - Project page: [JoyAI-Echo 1.5](https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/) - Paper: [Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds](https://arxiv.org/abs/2608.23383) - Inference code: [`echo_longvideo`](https://github.com/jd-opensource/JoyAI-Echo/tree/main/echo_longvideo) - Director Agent: [`Director_Agent`](https://github.com/jd-opensource/JoyAI-Echo/tree/main/echo_longvideo/Director_Agent) ## Acknowledgements We gratefully acknowledge the open-source projects that make this release possible, especially [LTX-2.3](https://huggingface.co/Lightricks/LTX-2.3), [Gemma](https://huggingface.co/google/gemma-3-12b-it), and [MSST-WebUI](https://github.com/SUC-DriverOld/MSST-WebUI). ## Citation ```bibtex @article{duan2026joyaiecho15, title = {Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds}, author = {Duan, Nan and Huang, Haoyang and Jin, Weiyang and Li, Haoran and Li, Yaowei and Li, Yuming and Liu, Yijun and Lu, Xin and Ma, Xiaoxiao and Ma, Yanwen and Su, Yaofeng and Sun, Yilang and Wang, Haoyu and Xue, Zeyue and Zhang, Songchun and Zhuang, Junhao}, journal = {arXiv preprint arXiv:2608.23383}, year = {2026}, eprint = {2608.23383}, archivePrefix = {arXiv}, primaryClass = {cs.CV}, url = {https://arxiv.org/abs/2608.23383} } ``` ## License This project is based on LTX-2 by Lightricks Ltd. All original copyright, license, patent, trademark, and attribution notices from LTX-2 are retained. This project remains subject to the LTX-2 Community License Agreement.