English
Related papers

Related papers: HuMMan: Multi-Modal 4D Human Dataset for Versatile…

200 papers

Synthetic data has emerged as a promising source for 3D human research as it offers low-cost access to large-scale human datasets. To advance the diversity and annotation quality of human models, we introduce a new synthetic dataset,…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Zhitao Yang , Zhongang Cai , Haiyi Mei , Shuai Liu , Zhaoxi Chen , Weiye Xiao , Yukun Wei , Zhongfei Qing , Chen Wei , Bo Dai , Wayne Wu , Chen Qian , Dahua Lin , Ziwei Liu , Lei Yang

Human following is a crucial feature of human-robot interaction, yet it poses numerous challenges to mobile agents in real-world scenarios. Some major hurdles are that the target person may be in a crowd, obstructed by others, or facing…

Robotics · Computer Science 2023-09-25 Mario Srouji , Yao-Hung Hubert Tsai , Hugues Thomas , Jian Zhang

While reconstructing human poses in 3D from inexpensive sensors has advanced significantly in recent years, quantifying the dynamics of human motion, including the muscle-generated joint torques and external forces, remains a challenge.…

Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous manipulation of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Xintao Lv , Liang Xu , Yichao Yan , Xin Jin , Congsheng Xu , Shuwen Wu , Yifan Liu , Lincheng Li , Mengxiao Bi , Wenjun Zeng , Xiaokang Yang

Existing text-to-image models still struggle to generate images of multiple objects, especially in handling their spatial positions, relative sizes, overlapping, and attribute bindings. To efficiently address these challenges, we develop a…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Sen Li , Ruochen Wang , Cho-Jui Hsieh , Minhao Cheng , Tianyi Zhou

We present Multi-HMR, a strong sigle-shot model for multi-person 3D human mesh recovery from a single RGB image. Predictions encompass the whole body, i.e., including hands and facial expressions, using the SMPL-X parametric model and 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Fabien Baradel , Matthieu Armando , Salma Galaaoui , Romain Brégier , Philippe Weinzaepfel , Grégory Rogez , Thomas Lucas

Human-centric scene understanding is significant for real-world applications, but it is extremely challenging due to the existence of diverse human poses and actions, complex human-environment interactions, severe occlusions in crowds, etc.…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Yiteng Xu , Peishan Cong , Yichen Yao , Runnan Chen , Yuenan Hou , Xinge Zhu , Xuming He , Jingyi Yu , Yuexin Ma

Representing human performance at high-fidelity is an essential building block in diverse applications, such as film production, computer games or videoconferencing. To close the gap to production-level quality, we introduce HumanRF, a 4D…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Mustafa Işık , Martin Rünz , Markos Georgopoulos , Taras Khakhulin , Jonathan Starck , Lourdes Agapito , Matthias Nießner

Recent years have witnessed tremendous progress in the 3D reconstruction of dynamic humans from a monocular video with the advent of neural rendering techniques. This task has a wide range of applications, including the creation of virtual…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Kanghao Chen , Zeyu Wang , Lin Wang

MEx: Multi-modal Exercises Dataset is a multi-sensor, multi-modal dataset, implemented to benchmark Human Activity Recognition(HAR) and Multi-modal Fusion algorithms. Collection of this dataset was inspired by the need for recognising and…

Computer Vision and Pattern Recognition · Computer Science 2019-08-27 Anjana Wijekoon , Nirmalie Wiratunga , Kay Cooper

We present HOI4D, a large-scale 4D egocentric dataset with rich annotations, to catalyze the research of category-level human-object interaction. HOI4D consists of 2.4M RGB-D egocentric video frames over 4000 sequences collected by 4…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Yunze Liu , Yun Liu , Che Jiang , Kangbo Lyu , Weikang Wan , Hao Shen , Boqiang Liang , Zhoujie Fu , He Wang , Li Yi

Photorealistic avatars of human faces have come a long way in recent years, yet research along this area is limited by a lack of publicly available, high-quality datasets covering both, dense multi-view camera captures, and rich facial…

Manipulating objects without grasping them is an essential component of human dexterity, referred to as non-prehensile manipulation. Non-prehensile manipulation may enable more complex interactions with the objects, but also presents…

Robotics · Computer Science 2024-07-16 Wenxuan Zhou , Bowen Jiang , Fan Yang , Chris Paxton , David Held

Human vision is capable of transforming two-dimensional observations into an egocentric three-dimensional scene understanding, which underpins the ability to translate complex scenes and exhibit adaptive behaviors. This capability, however,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Pei Liu , Hongliang Lu , Haichao Liu , Haipeng Liu , Xin Liu , Ruoyu Yao , Shengbo Eben Li , Jun Ma

Current approaches for humanoid whole-body manipulation, primarily relying on teleoperation or visual sim-to-real reinforcement learning, are hindered by hardware logistics and complex reward engineering. Consequently, demonstrated…

We present HumanEdit, a high-quality, human-rewarded dataset specifically designed for instruction-guided image editing, enabling precise and diverse image manipulations through open-form language instructions. Previous large-scale editing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Jinbin Bai , Wei Chow , Ling Yang , Xiangtai Li , Juncheng Li , Hanwang Zhang , Shuicheng Yan

Can Multimodal Large Language Models (MLLMs) develop an intuitive number sense similar to humans? Targeting this problem, we introduce Visual Number Benchmark (VisNumBench) to evaluate the number sense abilities of MLLMs across a wide range…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Tengjin Weng , Jingyi Wang , Wenhao Jiang , Zhong Ming

Human activity recognition (HAR) based on multimodal sensors has become a rapidly growing branch of biometric recognition and artificial intelligence. However, how to fully mine multimodal time series data and effectively learn accurate…

Computer Vision and Pattern Recognition · Computer Science 2022-05-25 Jialiang Wang , Haotian Wei , Yi Wang , Shu Yang , Chi Li

Human Action Recognition (HAR) is a very crucial task in computer vision. It helps to carry out a series of downstream tasks, like understanding human behaviors. Due to the complexity of human behaviors, many highly valuable behaviors are…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Hongwu Li , Zhenliang Zhang , Wei Wang

Biomedical data is inherently multimodal, consisting of electronic health records, medical imaging, digital pathology, genome sequencing, wearable sensors, and more. The application of artificial intelligence tools to these multifaceted…

Machine Learning · Computer Science 2024-08-26 Shentong Mo , Paul Pu Liang
‹ Prev 1 4 5 6 7 8 10 Next ›