中文
相关论文

相关论文: OO-dMVMT: A Deep Multi-view Multi-task Classificat…

200 篇论文

3D human pose estimation is a key enabling technology for applications such as healthcare monitoring, human-robot collaboration, and immersive gaming, but real-world deployment remains challenged by viewpoint variations. Existing methods…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Yejia Liu , Hengle Jiang , Haoxian Liu , Runxi Huang , Xiaomin Ouyang

Effective scene representation is critical for the visual grounding ability of representations, yet existing methods for 3D Visual Grounding are often constrained. They either only focus on geometric and visual cues, or, like traditional 3D…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Qinghongbing Xie , Zijian Liang , Fuhao Li , Long Zeng

To autonomously navigate and plan interactions in real-world environments, robots require the ability to robustly perceive and map complex, unstructured surrounding scenes. Besides building an internal representation of the observed scene…

机器人学 · 计算机科学 2021-05-18 Margarita Grinvald , Fadri Furrer , Tonci Novkovic , Jen Jen Chung , Cesar Cadena , Roland Siegwart , Juan Nieto

Gestures are a key component of non-verbal communication in traffic, often helping pedestrian-to-driver interactions when formal traffic rules may be insufficient. This problem becomes more apparent when autonomous vehicles (AVs) struggle…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Alif Rizqullah Mahdi , Mahdi Rezaei , Natasha Merat

LiDAR is crucial for robust 3D scene perception in autonomous driving. LiDAR perception has the largest body of literature after camera perception. However, multi-task learning across tasks like detection, segmentation, and motion…

计算机视觉与模式识别 · 计算机科学 2024-11-20 Sambit Mohapatra , Senthil Yogamani , Varun Ravi Kumar , Stefan Milz , Heinrich Gotzig , Patrick Mäder

Recent vision-language pre-trained models (VL-PTMs) have shown remarkable success in open-vocabulary tasks. However, downstream use cases often involve further fine-tuning of VL-PTMs, which may distort their general knowledge and impair…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Lin Zhu , Yifeng Yang , Qinying Gu , Xinbing Wang , Chenghu Zhou , Nanyang Ye

Hand detection is essential for many hand related tasks, e.g. parsing hand pose, understanding gesture, which are extremely useful for robotics and human-computer interaction. However, hand detection in uncontrolled environments is…

计算机视觉与模式识别 · 计算机科学 2016-12-09 Xiaoming Deng , Ye Yuan , Yinda Zhang , Ping Tan , Liang Chang , Shuo Yang , Hongan Wang

We present a new multi-stream 3D mesh reconstruction network (MSMR-Net) for hand pose estimation from a single RGB image. Our model consists of an image encoder followed by a mesh-convolution decoder composed of connected graph convolution…

计算机视觉与模式识别 · 计算机科学 2021-04-27 Uri Wollner , Guy Ben-Yosef

Tactile recognition of 3D objects remains a challenging task. Compared to 2D shapes, the complex geometry of 3D surfaces requires richer tactile signals, more dexterous actions, and more advanced encoding techniques. In this work, we…

计算机视觉与模式识别 · 计算机科学 2023-03-01 Jingxi Xu , Han Lin , Shuran Song , Matei Ciocarlie

Sensor-based human activity recognition is a key technology for many human-centered intelligent applications. However, this research is still in its infancy and faces many unresolved challenges. To address these, we propose a comprehensive…

信号处理 · 电气工程与系统科学 2025-04-08 Hanyu Liu , Ying Yu , Hang Xiao , Siyao Li , Xuze Li , Jiarui Li , Haotian Tang

Head gesture is a natural means of face-to-face communication between people but the recognition of head gestures in the context of virtual reality and use of head gesture as an interface for interacting with virtual avatars and virtual…

人机交互 · 计算机科学 2018-02-06 Jingbo Zhao , Robert S. Allison

We present a comprehensive framework for egocentric interaction recognition using markerless 3D annotations of two hands manipulating objects. To this end, we propose a method to create a unified dataset for egocentric 3D interaction…

计算机视觉与模式识别 · 计算机科学 2021-08-25 Taein Kwon , Bugra Tekin , Jan Stuhmer , Federica Bogo , Marc Pollefeys

Visual odometry networks commonly use pretrained optical flow networks in order to derive the ego-motion between consecutive frames. The features extracted by these networks represent the motion of all the pixels between frames. However,…

计算机视觉与模式识别 · 计算机科学 2020-11-18 Hamed Damirchi , Rooholla Khorrambakht , Hamid D. Taghirad

3D hand pose tracking/estimation will be very important in the next generation of human-computer interaction. Most of the currently available algorithms rely on low-cost active depth sensors. However, these sensors can be easily interfered…

计算机视觉与模式识别 · 计算机科学 2016-10-25 Jiawei Zhang , Jianbo Jiao , Mingliang Chen , Liangqiong Qu , Xiaobin Xu , Qingxiong Yang

Purpose: We describe a 3D multi-view perception system for the da Vinci surgical system to enable Operating room (OR) scene understanding and context awareness. Methods: Our proposed system is comprised of four Time-of-Flight (ToF) cameras…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Zhaoshuo Li , Amirreza Shaban , Jean-Gabriel Simard , Dinesh Rabindran , Simon DiMaio , Omid Mohareri

We present GestOS, a gesture-based operating system for high-level control of heterogeneous robot teams. Unlike prior systems that map gestures to fixed commands or single-agent actions, GestOS interprets hand gestures semantically and…

机器人学 · 计算机科学 2025-09-19 Artem Lykov , Oleg Kobzarev , Dzmitry Tsetserukou

Human-to-Robot handovers are useful for many Human-Robot Interaction scenarios. It is important to recognize when a human intends to initiate handovers, so that the robot does not try to take objects from humans when a handover is not…

计算机视觉与模式识别 · 计算机科学 2021-01-01 Jun Kwan , Chinkye Tan , Akansel Cosgun

Recent advances in machine learning technology have enabled highly portable and performant models for many common tasks, especially in image recognition. One emerging field, 3D human pose recognition extrapolated from video, has now…

人工智能 · 计算机科学 2022-03-24 Alex Moran , Bart Gebka , Joshua Goldshteyn , Autumn Beyer , Nathan Johnson , Alexander Neuwirth

Recovering high-fidelity 3D hand geometry from images is a critical task in computer vision, holding significant value for domains such as robotics, animation and VR/AR. Crucially, scalable applications demand both accuracy and deployment…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yumeng Liu , Xiao-Xiao Long , Marc Habermann , Xuanze Yang , Cheng Lin , Yuan Liu , Yuexin Ma , Wenping Wang , Ligang Liu

The gesture recognition using motion capture data and depth sensors has recently drawn more attention in vision recognition. Currently most systems only classify dataset with a couple of dozens different actions. Moreover, feature…

计算机视觉与模式识别 · 计算机科学 2014-09-02 Kyunghyun Cho , Xi Chen