SIRUS: Training-Free Concept Unlearning for Text-to-Video Models
Researchers have introduced SIRUS, a training-free, inference-time framework designed to suppress specific target concepts in text-to-video (T2V) generation models. SIRUS achieves 70.4% average forgetting success on the CogVideoX benchmark while minimizing video quality degradation, outperforming existing baselines such as VideoEraser. The work also presents a new video-centric evaluation framework for assessing T2V unlearning methods.
Why it matters: This approach enables safer and more controllable video generation by allowing unwanted concepts to be removed without retraining, addressing both practical and ethical concerns.
Full story at: arXiv Computer Vision ↗