中文
相关论文

相关论文: AnthroTAP: Learning Point Tracking with Real-World…

200 篇论文

Recent open-world representation learning approaches have leveraged CLIP to enable zero-shot 3D object recognition. However, performance on real point clouds with occlusions still falls short due to unrealistic pretraining settings.…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Khanh Nguyen , Ghulam Mubashar Hassan , Ajmal Mian

We tackle the challenges of synthesizing versatile, physically simulated human motions for full-body object manipulation. Unlike prior methods that are focused on detailed motion tracking, trajectory following, or teleoperation, our…

机器人学 · 计算机科学 2025-12-12 Chen Tessler , Yifeng Jiang , Erwin Coumans , Zhengyi Luo , Gal Chechik , Xue Bin Peng

Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, learning physical interaction remains constrained by the lack of large, diverse, and richly…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Yufan Deng , Daquan Zhou

We present MAMMA, a markerless motion-capture pipeline that accurately recovers SMPL-X parameters from multi-view video of two-person interaction sequences. Traditional motion-capture systems rely on physical markers. Although they offer…

Humans possess the cognitive ability to comprehend scenes in a compositional manner. To empower AI systems with similar capabilities, object-centric learning aims to acquire representations of individual objects from visual scenes without…

计算机视觉与模式识别 · 计算机科学 2023-09-07 Yinxuan Huang , Tonglin Chen , Zhimeng Shen , Jinghao Huang , Bin Li , Xiangyang Xue

Deep learning-based methods for video pedestrian detection and tracking require large volumes of training data to achieve good performance. However, data acquisition in crowded public environments raises data privacy concerns -- we are not…

Human performance capture is a highly important computer vision problem with many applications in movie production and virtual/augmented reality. Many previous performance capture approaches either required expensive multi-view setups or…

计算机视觉与模式识别 · 计算机科学 2021-11-23 Marc Habermann , Weipeng Xu , Michael Zollhoefer , Gerard Pons-Moll , Christian Theobalt

Point tracking aims to identify the same physical point across video frames and serves as a geometry-aware representation of motion. This representation supports a wide range of applications, from robotics to augmented reality, by enabling…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Görkay Aydemir

Human poses and motions are important cues for analysis of videos with people and there is strong evidence that representations based on body pose are highly effective for a variety of tasks such as activity recognition, content retrieval…

计算机视觉与模式识别 · 计算机科学 2018-04-12 Mykhaylo Andriluka , Umar Iqbal , Eldar Insafutdinov , Leonid Pishchulin , Anton Milan , Juergen Gall , Bernt Schiele

Humans naturally integrate vision and haptics for robust object perception during manipulation. The loss of either modality significantly degrades performance. Inspired by this multisensory integration, prior object pose estimation research…

机器人学 · 计算机科学 2025-09-12 Hongyu Li , Mingxi Jia , Tuluhan Akbulut , Yu Xiang , George Konidaris , Srinath Sridhar

Understanding human behaviour in crowded indoor environments is central to surveillance, smart buildings, and human-robot interaction, yet existing datasets rarely capture real-world indoor complexity at scale. We introduce IndoorCrowd, a…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Sebastian-Ion Nae , Radu Moldoveanu , Alexandra Stefania Ghita , Adina Magda Florea

Large datasets are the cornerstone of recent advances in computer vision using deep learning. In contrast, existing human motion capture (mocap) datasets are small and the motions limited, hampering progress on learning models of human…

计算机视觉与模式识别 · 计算机科学 2019-04-09 Naureen Mahmood , Nima Ghorbani , Nikolaus F. Troje , Gerard Pons-Moll , Michael J. Black

For multi-target tracking, target representation plays a crucial rule in performance. State-of-the-art approaches rely on the deep learning-based visual representation that gives an optimal performance at the cost of high computational…

计算机视觉与模式识别 · 计算机科学 2020-06-12 Mohib Ullah , Maqsood Mahmud , Habib Ullah , Kashif Ahmad , Ali Shariq Imran , Faouzi Alaya Cheikh

Predicting future trajectories for other road agents is an essential task for autonomous vehicles. Established trajectory prediction methods primarily use agent tracks generated by a detection and tracking system and HD map as inputs. In…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Seokha Moon , Hyun Woo , Hongbeen Park , Haeji Jung , Reza Mahjourian , Hyung-gun Chi , Hyerin Lim , Sangpil Kim , Jinkyu Kim

Humanoid control is an important research challenge offering avenues for integration into human-centric infrastructures and enabling physics-driven humanoid animations. The daunting challenges in this field stem from the difficulty of…

In the animation industry, cartoon videos are usually produced at low frame rate since hand drawing of such frames is costly and time-consuming. Therefore, it is desirable to develop computational models that can automatically interpolate…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Li Siyao , Shiyu Zhao , Weijiang Yu , Wenxiu Sun , Dimitris N. Metaxas , Chen Change Loy , Ziwei Liu

While deep-learning based tracking methods have achieved substantial progress, they entail large-scale and high-quality annotated data for sufficient training. To eliminate expensive and exhaustive annotation, we study self-supervised…

计算机视觉与模式识别 · 计算机科学 2023-01-02 Xin Li , Wenjie Pei , Yaowei Wang , Zhenyu He , Huchuan Lu , Ming-Hsuan Yang

This paper addresses the problem of 3D human pose estimation in the wild. A significant challenge is the lack of training data, i.e., 2D images of humans annotated with 3D poses. Such data is necessary to train state-of-the-art CNN…

计算机视觉与模式识别 · 计算机科学 2016-10-31 Grégory Rogez , Cordelia Schmid

Traditional temporal action detection (TAD) usually handles untrimmed videos with small number of action instances from a single label (e.g., ActivityNet, THUMOS). However, this setting might be unrealistic as different classes of actions…

计算机视觉与模式识别 · 计算机科学 2023-03-22 Jing Tan , Xiaotong Zhao , Xintian Shi , Bin Kang , Limin Wang

Humans make extensive use of vision and touch as complementary senses, with vision providing global information about the scene and touch measuring local information during manipulation without suffering from occlusions. While prior work…

机器人学 · 计算机科学 2023-08-01 Justin Kerr , Huang Huang , Albert Wilcox , Ryan Hoque , Jeffrey Ichnowski , Roberto Calandra , Ken Goldberg