中文
相关论文

相关论文: ExpertEdit: Learning Skill-Aware Motion Editing fr…

200 篇论文

Due to the challenges of manually collecting accurate editing data, existing datasets are typically constructed using various automated methods, leading to noisy supervision signals caused by the mismatch between editing instructions and…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Ming Li , Xin Gu , Fan Chen , Xiaoying Xing , Longyin Wen , Chen Chen , Sijie Zhu

As the world changes, we need to be able to update our models and correct false information without costly retraining. Knowledge-based model editing enables precise modifications to the weights of large language models in order to modify…

Automated tools for video editing and assembly have applications ranging from filmmaking and advertisement to content creation for social media. Previous video editing work has mainly focused on either retrieval or user interfaces, leaving…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Marcelo Sandoval-Castaneda , Bryan Russell , Josef Sivic , Gregory Shakhnarovich , Fabian Caba Heilbron

Video editing using diffusion models has achieved remarkable results in generating high-quality edits for videos. However, current methods often rely on large-scale pretraining, limiting flexibility for specific edits. First-frame-guided…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Chenjian Gao , Lihe Ding , Xin Cai , Zhanpeng Huang , Zibin Wang , Tianfan Xue

Recent image editing models boast next-level intelligent capabilities, facilitating cognition- and creativity-informed image editing. Yet, existing benchmarks provide too narrow a scope for evaluation, failing to holistically assess these…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Kaihang Pan , Weile Chen , Haiyi Qiu , Qifan Yu , Wendong Bu , Zehan Wang , Yun Zhu , Juncheng Li , Siliang Tang

We present EgoExo-Fitness, a new full-body action understanding dataset, featuring fitness sequence videos recorded from synchronized egocentric and fixed exocentric (third-person) cameras. Compared with existing full-body action…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Yuan-Ming Li , Wei-Jin Huang , An-Lan Wang , Ling-An Zeng , Jing-Ke Meng , Wei-Shi Zheng

Despite recent advances in inversion and instruction-based image editing, existing approaches primarily excel at editing single, prominent objects but significantly struggle when applied to complex scenes containing multiple entities. To…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Bimsara Pathiraja , Maitreya Patel , Shivam Singh , Yezhou Yang , Chitta Baral

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric)…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Haoyu Zhang , Qiaohui Chu , Meng Liu , Haoxiang Shi , Yaowei Wang , Liqiang Nie

Timely and transparent feedback is essential for effective surgical training, yet current assessment remains dependent on expert observation, limiting scalability and opportunities for autonomous practice. We present ExpOS, an explainable…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Roi Papo , Idan Smoller , Shlomi Laufer

Pre-trained Vision Transformers (ViTs) are increasingly deployed for medical image classification. However, correcting their inevitable failure cases in dynamic clinical scenarios poses a critical challenge. Conventional fine-tuning…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Yuanye Liu , Siyuan Zhou , Ke Zhang , Lei Li , Wei Chen , Xiahai Zhuang

Motion instruction is a crucial task that helps athletes refine their technique by analyzing movements and providing corrective guidance. Although recent advances in multimodal models have improved motion understanding, generating precise…

计算与语言 · 计算机科学 2025-09-16 Wei-Hsin Yeh , Yu-An Su , Chih-Ning Chen , Yi-Hsueh Lin , Calvin Ku , Wen-Hsin Chiu , Min-Chun Hu , Lun-Wei Ku

We propose ClipFace, a novel self-supervised approach for text-guided editing of textured 3D morphable model of faces. Specifically, we employ user-friendly language prompts to enable control of the expressions as well as appearance of 3D…

计算机视觉与模式识别 · 计算机科学 2023-04-25 Shivangi Aneja , Justus Thies , Angela Dai , Matthias Nießner

Many application areas ranging from serious games for health to learning by demonstration in robotics, could benefit from large body movement datasets extracted from textual instructions accompanied by images. The interpretation of…

人机交互 · 计算机科学 2020-06-09 Himangshu Sarma , Robert Porzel , Jan Smeddinck , Rainer Malaka

We introduce an approach for augmenting text-to-video generation models with customized motions, extending their capabilities beyond the motions depicted in the original training data. By leveraging a few video samples demonstrating…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Joanna Materzynska , Josef Sivic , Eli Shechtman , Antonio Torralba , Richard Zhang , Bryan Russell

Recent advances in text-guided image editing enable users to perform image edits through simple text inputs, leveraging the extensive priors of multi-step diffusion-based text-to-image models. However, these methods often fall short of the…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Trong-Tung Nguyen , Quang Nguyen , Khoi Nguyen , Anh Tran , Cuong Pham

Combining Vision Large Language Models (VLLMs) with diffusion models offers a powerful method for executing image editing tasks based on human language instructions. However, language instructions alone often fall short in accurately…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Tianshuo Yuan , Yuxiang Lin , Jue Wang , Zhi-Qi Cheng , Xiaolong Wang , Jiao GH , Wei Chen , Xiaojiang Peng

3D object editing is essential for interactive content creation in gaming, animation, and robotics, yet current approaches remain inefficient, inconsistent, and often fail to preserve unedited regions. Most methods rely on editing…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Junliang Ye , Shenghao Xie , Ruowen Zhao , Zhengyi Wang , Hongyu Yan , Wenqiang Zu , Lei Ma , Jun Zhu

Understanding how images of objects and scenes behave in response to specific ego-motions is a crucial aspect of proper visual development, yet existing visual learning methods are conspicuously disconnected from the physical source of…

计算机视觉与模式识别 · 计算机科学 2016-03-30 Dinesh Jayaraman , Kristen Grauman

Inversion-based visual editing provides an effective and training-free way to edit an image or a video based on user instructions. Existing methods typically inject source image information during the sampling process to maintain editing…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Zhi Ouyang , Dian Zheng , Xiao-Ming Wu , Jian-Jian Jiang , Kun-Yu Lin , Jingke Meng , Wei-Shi Zheng

Visually-guided image editing, where edits are conditioned on both visual cues and textual prompts, has emerged as a powerful paradigm for fine-grained, controllable content generation. Although recent generative models have shown…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Sara Ghazanfari , Wei-An Lin , Haitong Tian , Ersin Yumer