CoinVE-Edit — Compositional Instruction-Guided Video Editing

Apply several edits at once to one video. CoinVE-Edit grounds each instruction to its own spatio-temporal region with a predicted mask, then injects the per-instruction features through a region-aware residual attention branch on a Wan2.1-T2V-14B DiT, so the edits compose instead of interfering with each other.

Paper · Code · Model

17 49
Resolution
15 50
0 2147483647
Examples