中文用于鲁棒视频编辑的遮挡感知物理语义关键帧选择
ENOcclusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing
针对视频编辑在遮挡、视角变化和快速运动下的不稳定性,本文提出利用3D重建作为视觉锚点,提升编辑一致性。方法先重建场景3D表示,再将编辑指令传播到各帧,有效解决定位不准和闪烁问题。实际意义:显著增强复杂场景下的视频编辑鲁棒性与可操作性。
arXiv:2605.23192v1 Announce Type: new Abstract: Video editing has recently achieved remarkable progress with diffusion-based generative models, enabling diverse object-level manipulations from natural language instructions. However, existing methods often struggle under occlusion, viewpoint changes, and fast object motion, where unreliable visual observations lead to inaccurate localization, temporal flickering, and inconsistent edits. In this work, we identify the absence of reliable visual anchors as a fundamental bottleneck in occlusion-robust video editing. To address this issue, we propos