← Back to brief
ResearchOfficialPreprintarXiv Computer Vision

Moving Alphabet: Controlled Study Reveals How Data Distribution and Caption Quality Impact Text-to-Video Generation

A new preprint introduces Moving Alphabet, a procedural testbed designed for controlled experiments on text-to-video training data. The study finds that a diverse and balanced distribution of video content and duration is critical for model generalization, and that caption quality significantly affects both performance and training efficiency. It also shows that classifier-free guidance and fine-tuning on high-quality data can only partially recover performance if pre-training data is poor.

Why it matters: This research provides systematic insights into the underexplored role of training data quality and distribution in text-to-video models, offering guidance for effective data curation.

Full story at: arXiv Computer Vision