中文
相关论文

相关论文: MoCrop: Training Free Motion Guided Cropping for E…

200 篇论文

Monocular 3D motion capture (mocap) is beneficial to many applications. The use of a single camera, however, often fails to handle occlusions of different body parts and hence it is limited to capture relatively simple movements. We present…

计算机视觉与模式识别 · 计算机科学 2023-03-07 Han Liang , Yannan He , Chengfeng Zhao , Mutian Li , Jingya Wang , Jingyi Yu , Lan Xu

Recently, Transformer networks have achieved impressive results on a variety of vision tasks. However, most of them are computationally expensive and not suitable for real-world mobile applications. In this work, we present Mobile…

计算机视觉与模式识别 · 计算机科学 2022-05-27 Hailong Ma , Xin Xia , Xing Wang , Xuefeng Xiao , Jiashi Li , Min Zheng

Human motion recovery for real-world interaction demands both precise action details and metric-scale trajectories. Recovering absolute human pose from monocular input presents a viable solution, but faces two main challenges: (1) models'…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Zhumei Wang , Zechen Hu , Ruoxi Guo , Huaijin Pi , Ziyong Feng , Liang Zhang , Mingtao Pei , Siyuan Huang

Adopting contrastive image-text pretrained models like CLIP towards video classification has gained attention due to its cost-effectiveness and competitive performance. However, recent works in this area face a trade-off. Finetuning the…

计算机视觉与模式识别 · 计算机科学 2023-04-10 Syed Talal Wasim , Muzammal Naseer , Salman Khan , Fahad Shahbaz Khan , Mubarak Shah

Cross-modal pre-training has shown impressive performance on a wide range of downstream tasks, benefiting from massive image-text pairs collected from the Internet. In practice, online data are growing constantly, highlighting the…

计算机视觉与模式识别 · 计算机科学 2023-08-23 Xinchi Deng , Han Shi , Runhui Huang , Changlin Li , Hang Xu , Jianhua Han , James Kwok , Shen Zhao , Wei Zhang , Xiaodan Liang

Image cropping is crucial for enhancing the visual appeal and narrative impact of photographs, yet existing rule-based and data-driven approaches often lack diversity or require annotated training data. We introduce ProCrop, a…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Ke Zhang , Tianyu Ding , Jiachen Jiang , Tianyi Chen , Ilya Zharkov , Vishal M. Patel , Luming Liang

Markerless human motion capture (mocap) from multiple RGB cameras is a widely studied problem. Existing methods either need calibrated cameras or calibrate them relative to a static camera, which acts as the reference frame for the mocap…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Nitin Saini , Chun-hao P. Huang , Michael J. Black , Aamir Ahmad

Motion capture from a monocular video is fundamental and crucial for us humans to naturally experience and interact with each other in Virtual Reality (VR) and Augmented Reality (AR). However, existing methods still struggle with…

计算机视觉与模式识别 · 计算机科学 2022-10-31 Xin Chen , Zhuo Su , Lingbo Yang , Pei Cheng , Lan Xu , Bin Fu , Gang Yu

Cluttered bin-picking environments are challenging for pose estimation models. Despite the impressive progress enabled by deep learning, single-view RGB pose estimation models perform poorly in cluttered dynamic environments. Imbuing the…

机器人学 · 计算机科学 2026-02-02 Arul Selvam Periyasamy , Sven Behnke

Existing human Motion Capture (MoCap) methods mostly focus on the visual similarity while neglecting the physical plausibility. As a result, downstream tasks such as driving virtual human in 3D scene or humanoid robots in real world suffer…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Shenghao Ren , Yi Lu , Jiayi Huang , Jiayi Zhao , He Zhang , Tao Yu , Qiu Shen , Xun Cao

The primary challenge in accelerating image super-resolution lies in reducing computation while maintaining performance and adaptability. Motivated by the observation that high-frequency regions (e.g., edges and textures) are most critical…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Wei Shang , Dongwei Ren , Wanying Zhang , Pengfei Zhu , Qinghua Hu , Wangmeng Zuo

Motion capture through tracking retroreflectors obtains highly accurate pose estimation, which is frequently used in robotics. Unlike commercial motion capture systems, fiducial marker-based tracking methods, such as AprilTags, can perform…

机器人学 · 计算机科学 2023-07-03 Gary Lvov , Mark Zolotas , Nathaniel Hanson , Austin Allison , Xavier Hubbard , Michael Carvajal , Taskin Padir

Learning real-world dynamics from visual observations is crucial for various domains. A common strategy is to calibrate simulators by estimating physical parameters, yet accuracy is ultimately bounded by the underlying physical models,…

机器学习 · 计算机科学 2026-05-22 Jiaxu Wang , Junhao He , Jingkai Sun , Yi Gu , Yunyang Mo , Jiahang Cao , Qiang Zhang , Renjing Xu

With the rapid proliferation of the Internet of Things, video analytics has become a cornerstone application in wireless multimedia sensor networks. To support such applications under bandwidth constraints, learning-based adaptive…

Most action recognition solutions rely on dense sampling to precisely cover the informative temporal clip. Extensively searching temporal region is expensive for a real-world application. In this work, we focus on improving the inference…

计算机视觉与模式识别 · 计算机科学 2021-07-30 Chunhui Liu , Xinyu Li , Hao Chen , Davide Modolo , Joseph Tighe

Drug Mechanism of Action (MoA) mainly investigates how drug molecules interact with cells, which is crucial for drug discovery and clinical application. Recently, deep learning models have been used to recognize MoA by relying on…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Fengqian Pang , Chunyue Lei , Hongfei Zhao , Chenghao Liu , Zhiqiang Xing , Huafeng Wang , Chuyang Ye

Autonomous motion capture (mocap) systems for outdoor scenarios involving flying or mobile cameras rely on i) a robotic front-end to track and follow a human subject in real-time while he/she performs physical activities, and ii) an…

In this paper, a marker-based, single-person optical motion capture method (DeepMoCap) is proposed using multiple spatio-temporally aligned infrared-depth sensors and retro-reflective straps and patches (reflectors). DeepMoCap explores…

计算机视觉与模式识别 · 计算机科学 2021-10-15 Anargyros Chatzitofis , Dimitrios Zarpalas , Stefanos Kollias , Petros Daras

We present Mobile Video Networks (MoViNets), a family of computation and memory efficient video networks that can operate on streaming video for online inference. 3D convolutional neural networks (CNNs) are accurate at video recognition but…

计算机视觉与模式识别 · 计算机科学 2021-04-20 Dan Kondratyuk , Liangzhe Yuan , Yandong Li , Li Zhang , Mingxing Tan , Matthew Brown , Boqing Gong

Animating images with interactive motion control has garnered popularity for image-to-video (I2V) generation. Modern approaches typically rely on large Gaussian kernels to extend motion trajectories as condition without explicitly defining…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Zhongwei Zhang , Fuchen Long , Zhaofan Qiu , Yingwei Pan , Wu Liu , Ting Yao , Tao Mei