Seedance 2.0
ByteDance·
Videomultimodalcloud
Overview
ByteDance text/image/audio-to-video model with unified audio-video joint generation. Generates 4-15s clips up to 1080p across multiple aspect ratios. Available via Volcano Engine and API.
Capabilities and innovations
Text-to-videoImage-to-videoAudio-video joint generationUp to 1080p output4-15s clipsUnified multimodal audio-video architectureMultimodal reference inputs (text/image/audio/video)Director-level camera & lighting control
Benchmarks
No tracked academic benchmark applies to video models; the timeline carries no scores and no leadership tracking for them.
- GPQA Diamond: not reported by the vendor
- Humanity’s Last Exam (no tools): not reported by the vendor
- SWE-bench Verified: not reported by the vendor
- SWE-bench Pro: not reported by the vendor
- MMLU: not reported by the vendor
- MMLU-Pro: not reported by the vendor
- Terminal-Bench 2.x: not reported by the vendor
- MMMU-Pro: not reported by the vendor
Links
More from ByteDance
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.