
JoyAI-Echo 1.5 is a long-form audio-video generation model from JD's Explore Academy, built for studios and e-commerce teams that need minutes of coherent, synchronized video with sound. Shown at the JD Discovery conference on September 9, 2026, it arrived alongside EchoWM, an open-sourced interactive audio-visual world model.
Core Features
- Generates more than 10 minutes of continuous audio-video while keeping character, voice, and story consistent.
- A Memory Bank architecture that holds multi-shot, long-horizon coherence across a full scene.
- Up to 7.5x inference speedup from long-video post-training, reaching about 24 frames per second at 480P on two H200 GPUs.
- A Director Agent that expands a natural-language brief into shot plans, reviews generations, and assembles the final clip.
- EchoWM jointly produces 720P video with ambient sound, music, and speech, supporting first-person navigation and multi-round exploration.
Use Cases
- E-commerce brands creating product and advertising videos from a single brief.
- Short-drama and AI comic teams needing stable characters across episodes.
- Embodied-AI researchers simulating environments with native audiovisual output.
- Creators exploring interactive worlds they can enter and steer.
Pricing
JoyAI-Echo 1.5 was presented as a JD research release, with EchoWM open-sourced under a public license; JD states broader availability arrives in October 2026. Concrete API pricing was not disclosed at launch, so teams should treat access as preview-stage and confirm commercial terms through JD's AI cloud offerings when they open.
Our Take
Best for long-form, story-driven video where most models break coherence after a few seconds. The trade-off is a preview-era status with no public price yet, so budget production on it only after JD publishes commercial terms. Explore more AI Video tools.










