中文
相关论文

相关论文: RNA: Video Editing with ROI-based Neural Atlas

200 篇论文

Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a second, drag editing still lags behind due to the challenge of…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Joonghyuk Shin , Daehyeon Choi , Jaesik Park

In NeRF-aided editing tasks, object movement presents difficulties in supervision generation due to the introduction of variability in object positions. Moreover, the removal operations of certain scene objects often lead to empty regions,…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Zhenyang Li , Zilong Chen , Feifan Qu , Mingqing Wang , Yizhou Zhao , Kai Zhang , Yifan Peng

Rapid growth in the development of medical imaging analysis technology has been propelled by the great interest in improving computer-aided diagnosis and detection (CAD) systems for three popular image visualization tasks: classification,…

计算机视觉与模式识别 · 计算机科学 2022-03-03 Manu Goyal , Moi Hoon Yap , Saeed Hassanpour

The ability of Generative Adversarial Networks to encode rich semantics within their latent space has been widely adopted for facial image editing. However, replicating their success with videos has proven challenging. Sets of high-quality…

计算机视觉与模式识别 · 计算机科学 2022-01-24 Rotem Tzaban , Ron Mokady , Rinon Gal , Amit H. Bermano , Daniel Cohen-Or

The recent advances in deep learning have made it possible to generate photo-realistic images by using neural networks and even to extrapolate video frames from an input video clip. In this paper, for the sake of both furthering this…

计算机视觉与模式识别 · 计算机科学 2018-08-10 Lijie Fan , Wenbing Huang , Chuang Gan , Junzhou Huang , Boqing Gong

Video generation models have shown their superior ability to generate photo-realistic video. However, how to accurately control (or edit) the video remains a formidable challenge. The main issues are: 1) how to perform direct and accurate…

图形学 · 计算机科学 2024-07-23 Yufan Deng , Ruida Wang , Yuhao Zhang , Yu-Wing Tai , Chi-Keung Tang

To enhance on-road environmental perception for autonomous driving, accurate and real-time analytics on high-resolution video frames generated from on-board cameras be-comes crucial. In this paper, we design a lightweight object location…

多媒体 · 计算机科学 2023-09-01 Yan Cheng , Peng Yang , Ning Zhang , Jiawei Hou

Despite the rapid growth in datasets for video activity, stable robust activity recognition with neural networks remains challenging. This is in large part due to the explosion of possible variation in video -- including lighting changes,…

计算机视觉与模式识别 · 计算机科学 2019-12-04 Yi Zhang , Xinyue Wei , Weichao Qiu , Zihao Xiao , Gregory D. Hager , Alan Yuille

Resembling the rapid learning capability of human, few-shot learning empowers vision systems to understand new concepts by training with few samples. Leading approaches derived from meta-learning on images with a single visual object.…

计算机视觉与模式识别 · 计算机科学 2020-03-17 Xiaopeng Yan , Ziliang Chen , Anni Xu , Xiaoxi Wang , Xiaodan Liang , Liang Lin

Robotic-assisted surgery (RAS) is established in clinical practice, and automated surgical skill assessment utilizing multimodal data offers transformative potential for surgical analytics and education. However, developing effective…

Recent advances in image editing, driven by image diffusion models, have shown remarkable progress. However, significant challenges remain, as these models often struggle to follow complex edit instructions accurately and frequently…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Noam Rotstein , Gal Yona , Daniel Silver , Roy Velich , David Bensaïd , Ron Kimmel

Recent advances in generative AI have significantly enhanced image and video editing, particularly in the context of text prompt control. State-of-the-art approaches predominantly rely on diffusion models to accomplish these tasks. However,…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Haoyu Ma , Shahin Mahdizadehaghdam , Bichen Wu , Zhipeng Fan , Yuchao Gu , Wenliang Zhao , Lior Shapira , Xiaohui Xie

With the popularity of implicit neural representations, or neural radiance fields (NeRF), there is a pressing need for editing methods to interact with the implicit 3D models for tasks like post-processing reconstructed scenes and 3D…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Xiangyu Wang , Jingsen Zhu , Qi Ye , Yuchi Huo , Yunlong Ran , Zhihua Zhong , Jiming Chen

NeRF's high-quality scene synthesis capability was quickly accepted by scholars in the years after it was proposed, and significant progress has been made in 3D scene representation and synthesis. However, the high computational cost limits…

计算机视觉与模式识别 · 计算机科学 2024-01-24 Shun Fang , Ming Cui , Xing Feng , Yanan Zhang

Rendering photorealistic and dynamically moving human heads is crucial for ensuring a pleasant and immersive experience in AR/VR and video conferencing applications. However, existing methods often struggle to model challenging facial…

计算机视觉与模式识别 · 计算机科学 2026-02-20 Cong Wang , Di Kang , Yan-Pei Cao , Linchao Bao , Ying Shan , Song-Hai Zhang

Deep convolutional neural networks have driven substantial advancements in the automatic understanding of images. Requiring a large collection of images and their associated annotations is one of the main bottlenecks limiting the adoption…

计算机视觉与模式识别 · 计算机科学 2019-08-22 Zahra Mirikharaji , Yiqi Yan , Ghassan Hamarneh

Advances in NERFs have allowed for 3D scene reconstructions and novel view synthesis. Yet, efficiently editing these representations while retaining photorealism is an emerging challenge. Recent methods face three primary limitations:…

Annotating long-horizon robotic demonstrations with precise temporal action boundaries is crucial for training and evaluating action segmentation and manipulation policy learning methods. Existing annotation tools, however, are often…

机器人学 · 计算机科学 2026-04-30 Sergej Stanovcic , Daniel Sliwowski , Dongheui Lee

Recent advancements in video generation highlight that realistic audio-visual synchronization is crucial for engaging content creation. However, existing video editing methods largely overlook audio-visual synchronization and lack the…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Haojie Zheng , Shuchen Weng , Jingqi Liu , Siqi Yang , Boxin Shi , Xinlong Wang

Face parsing computes pixel-wise label maps for different semantic components (e.g., hair, mouth, eyes) from face images. Existing face parsing literature have illustrated significant advantages by focusing on individual regions of interest…

计算机视觉与模式识别 · 计算机科学 2019-06-05 Jinpeng Lin , Hao Yang , Dong Chen , Ming Zeng , Fang Wen , Lu Yuan