中文
相关论文

相关论文: OWL: A Novel Approach to Machine Perception During…

200 篇论文

In this paper, we introduce a method to automatically reconstruct the 3D motion of a person interacting with an object from a single RGB video. Our method estimates the 3D poses of the person together with the object pose, the contact…

计算机视觉与模式识别 · 计算机科学 2021-11-03 Zongmian Li , Jiri Sedlar , Justin Carpentier , Ivan Laptev , Nicolas Mansard , Josef Sivic

WiFi-based sensing for human activity recognition (HAR) has recently become a hot topic as it brings great benefits when compared with video-based HAR, such as eliminating the demands of line-of-sight (LOS) and preserving privacy. Making…

信号处理 · 电气工程与系统科学 2022-06-22 Yanling Hao , Zhiyuan Shi , Yuanwei Liu

Open-world object detection (OWOD) is a challenging problem that combines object detection with incremental learning and open-set learning. Compared to standard object detection, the OWOD setting is task to: 1) detect objects seen during…

计算机视觉与模式识别 · 计算机科学 2023-02-24 Jinan Yu , Liyan Ma , Zhenglin Li , Yan Peng , Shaorong Xie

Efficient and high-accuracy 3D occupancy prediction is vital for the performance of autonomous driving systems. However, existing methods struggle to balance precision and efficiency: high-accuracy approaches are often hindered by heavy…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Yuchen Zhou , Yan Luo , Xiaogang Wang , Xingjian Gu , Mingzhou Lu , Xiangbo Shu

Ordered Weighted $L_{1}$ (OWL) regularized regression is a new regression analysis for high-dimensional sparse learning. Proximal gradient methods are used as standard approaches to solve OWL regression. However, it is still a burning issue…

机器学习 · 计算机科学 2021-10-20 Runxue Bao , Bin Gu , Heng Huang

We propose a novel method for learning convolutional neural image representations without manual supervision. We use motion cues in the form of optical flow, to supervise representations of static images. The obvious approach of training a…

计算机视觉与模式识别 · 计算机科学 2018-07-17 Aravindh Mahendran , James Thewlis , Andrea Vedaldi

Interpreting motion captured in image sequences is crucial for a wide range of computer vision applications. Typical estimation approaches include optical flow (OF), which approximates the apparent motion instantaneously in a scene, and…

计算机视觉与模式识别 · 计算机科学 2024-08-30 Tanner D. Harms , Steven L. Brunton , Beverley J. McKeon

We present a new trainable system for physically plausible markerless 3D human motion capture, which achieves state-of-the-art results in a broad range of challenging scenarios. Unlike most neural methods for human motion capture, our…

计算机视觉与模式识别 · 计算机科学 2021-05-04 Soshi Shimada , Vladislav Golyanik , Weipeng Xu , Patrick Pérez , Christian Theobalt

Multiple object tracking (MOT) has been successfully investigated in computer vision. However, MOT for the videos captured by unmanned aerial vehicles (UAV) is still challenging due to small object size, blurred object appearance, and very…

计算机视觉与模式识别 · 计算机科学 2023-08-16 Mufeng Yao , Jiaqi Wang , Jinlong Peng , Mingmin Chi , Chao Liu

While 3D object bounding box (bbox) representation has been widely used in autonomous driving perception, it lacks the ability to capture the precise details of an object's intrinsic geometry. Recently, occupancy has emerged as a promising…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Chaoda Zheng , Feng Wang , Naiyan Wang , Shuguang Cui , Zhen Li

Modeling and capturing the 3D spatial arrangement of the human and the object is the key to perceiving 3D human-object interaction from monocular images. In this work, we propose to use the Human-Object Offset between anchors which are…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Chaofan Huo , Ye Shi , Yuexin Ma , Lan Xu , Jingyi Yu , Jingya Wang

Looming, traditionally defined as the relative expansion of objects in the observer's retina, is a fundamental visual cue for perception of threat and can be used to accomplish collision free navigation. In this paper we derive novel…

图像与视频处理 · 电气工程与系统科学 2022-10-21 Juan Yepes , Daniel Raviv

Learning object-centric representations from complex natural environments enables both humans and machines with reasoning abilities from low-level perceptual features. To capture compositional entities of the scene, we proposed cyclic walks…

计算机视觉与模式识别 · 计算机科学 2023-11-02 Ziyu Wang , Mike Zheng Shou , Mengmi Zhang

We address the task of predicting what parts of an object can open and how they move when they do so. The input is a single image of an object, and as output we detect what parts of the object can open, and the motion parameters describing…

计算机视觉与模式识别 · 计算机科学 2022-03-31 Hanxiao Jiang , Yongsen Mao , Manolis Savva , Angel X. Chang

Traditional object detection models are constrained by the limitations of closed-set datasets, detecting only categories encountered during training. While multimodal models have extended category recognition by aligning text and image…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Lihao Liu , Juexiao Feng , Hui Chen , Ao Wang , Lin Song , Jungong Han , Guiguang Ding

We introduce Correspondence-Oriented Imitation Learning (COIL), a conditional policy learning framework for visuomotor control with a flexible task representation in 3D. At the core of our approach, each task is defined by the intended…

机器人学 · 计算机科学 2025-12-08 Yunhao Cao , Zubin Bhaumik , Jessie Jia , Xingyi He , Kuan Fang

Motion forecasting aims to predict the future trajectories of dynamic agents in the scene, enabling autonomous vehicles to effectively reason about scene evolution. Existing approaches operate under the closed-world regime and assume fixed…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Nicolas Schischka , Nikhil Gosala , B Ravi Kiran , Senthil Yogamani , Abhinav Valada

3D semantic occupancy prediction networks have demonstrated remarkable capabilities in reconstructing the geometric and semantic structure of 3D scenes, providing crucial information for robot navigation and autonomous driving systems.…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Junming Wang , Wei Yin , Xiaoxiao Long , Xingyu Zhang , Zebin Xing , Xiaoyang Guo , Qian Zhang

This paper addresses the challenge of robotic grasping of general objects. Similar to prior research, the task reads a single-view 3D observation (i.e., point clouds) captured by a depth camera as input. Crucially, the success of object…

机器人学 · 计算机科学 2024-07-23 Kangqi Ma , Hao Dong , Yadong Mu

Visual repetition is ubiquitous in our world. It appears in human activity (sports, cooking), animal behavior (a bee's waggle dance), natural phenomena (leaves in the wind) and in urban environments (flashing lights). Estimating visual…

计算机视觉与模式识别 · 计算机科学 2018-06-20 Tom F. H. Runia , Cees G. M. Snoek , Arnold W. M. Smeulders