In 2026, local video generation comes down to three open options: MiniMax H3 (joint audio-video generation, 15 s), Wan 2.2 (Alibaba Tongyi, the most mature ecosystem), and LTX-2.5 (extremely fast + native audio). They are not replacements for one another — they are three different trade-offs. Here's the verdict first:
| Model | Params | Local VRAM | Max duration | Native audio | License | Speed feel |
|---|---|---|---|---|---|---|
| MiniMax H3 | 33.1B | 8GB minimum / 16GB usable / 32GB comfortable | 15 s | ✅ Stereo | Community license (with region and revenue clauses) | Slow (5090 768p 5s 4-step ~80 s) |
| Wan 2.2 TI2V-5B | 5B | 8GB (GGUF) / 16GB comfortable | ~5 s | ❌ Needs separate audio | Apache 2.0 (most worry-free for commercial use) | Medium (4090 480p 5s ~4–6 min) |
| Wan 2.2 A14B | 27B (14B active) | 12GB (GGUF offload) / 40GB+ full quality | ~5 s | ❌ | Apache 2.0 | Slow (best-quality tier) |
| LTX-2.5 | 22B | 12GB (distilled/quantized) / 32GB+ full quality | ~10 s | ✅ Generated jointly in one pass | Open weights (check official terms for commercial use; tiered) | Very fast (distilled version close to real-time) |
| What matters most to you | Recommendation | Why |
|---|---|---|
| Characters talking, with sound effects generated in sync with the picture | H3 or LTX-2.5 | Both do native joint audio-video generation; Wan 2.2 needs a dubbing or lip-sync pass |
| A complete 10–15 s narrative clip | H3 | 15 s is the longest of the open models; LTX is ~10 s; Wan is ~5 s |
| Fast iteration and lots of drafts | LTX-2.5 distilled | Distillation plus near real-time speed on a single card — the fastest of the three |
| The most worry-free commercial license | Wan 2.2 | Apache 2.0; H3 and LTX have their own terms (H3 also adds region restrictions) |
| Ecosystem and number of LoRAs | Wan 2.2 | Out for over a year, with the richest training toolchain and community LoRAs; H3's ecosystem is filling in fast |
| Ceiling on image quality and detail | H3 (768p) and Wan A14B | Both are top-tier in their resolution class; LTX's strength is speed, not detail |
Note that "fast" comes in two flavors: LTX is fast thanks to architecture and distillation; H3's 4-step acceleration LoRA can also cut 20 steps to 4 (about 4×), but the base is different. If you mainly make short clips where audio must stay in sync, the speed gap between H3 and LTX can be leveled out by your workflow (draft resolution).