Happy Horse 1.0
Alibaba·
Videomultimodalcloud
Overview
Video generation model from Alibaba's ATH AI Innovation Unit. 15B-parameter single-stream Transformer (40 layers) with native audio/video/text token fusion. #1 in both text-to-video and image-to-video on Artificial Analysis Video Arena (Elo blind test). Initially released anonymously; confirmed as Alibaba on April 10, 2026. Beta — no public API or released weights.
Capabilities and innovations
Text-to-videoImage-to-videoNative audio/video/text token fusionSingle-stream 40-layer TransformerVideo Arena #1 (T2V and I2V)Native Multimodal Token Fusion
Benchmarks
No tracked academic benchmark applies to video models; the timeline carries no scores and no leadership tracking for them.
- GPQA Diamond: not reported by the vendor
- Humanity’s Last Exam (no tools): not reported by the vendor
- SWE-bench Verified: not reported by the vendor
- SWE-bench Pro: not reported by the vendor
- MMLU: not reported by the vendor
- MMLU-Pro: not reported by the vendor
- Terminal-Bench 2.x: not reported by the vendor
- MMMU-Pro: not reported by the vendor
Architecture and hardware
- Parameters
- 15B total, 15B active per token (dense)
Links
More from Alibaba
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.