中文
相关论文

相关论文: FLAG3D: A 3D Fitness Activity Dataset with Languag…

200 篇论文

The scarcity of high quality actions video data is a bottleneck in the research and application of action recognition. Although significant effort has been made in this area, there still exist gaps in the range of available data types a…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Shuo Wang , Amiya Ranjan , Lawrence Jiang

3D medical image analysis is pivotal in numerous clinical applications. However, the scarcity of labeled data and limited generalization capabilities hinder the advancement of AI-empowered models. Radiology reports are easily accessible and…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Xuefeng Ni , Linshan Wu , Jiaxin Zhuang , Qiong Wang , Mingxiang Wu , Varut Vardhanabhuti , Lihai Zhang , Hanyu Gao , Hao Chen

While multi-modality large language models excel in object-centric or indoor scenarios, scaling them to 3D city-scale environments remains a formidable challenge. To bridge this gap, we propose 3DCity-LLM, a unified framework designed for…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Yiping Chen , Jinpeng Li , Wenyu Ke , Yang Luo , Jie Ouyang , Zhongjie He , Li Liu , Hongchao Fan , Hao Wu

We introduce MotionScript, a novel framework for generating highly detailed, natural language descriptions of 3D human motions. Unlike existing motion datasets that rely on broad action labels or generic captions, MotionScript provides…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Payam Jome Yazdian , Rachel Lagasse , Hamid Mohammadi , Eric Liu , Li Cheng , Angelica Lim

3D Multi-modal Large Language Models (MLLMs) still lag behind their 2D peers, largely because large-scale, high-quality 3D scene-dialogue datasets remain scarce. Prior efforts hinge on expensive human annotation and leave two key…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Siyuan Wei , Chunjie Wang , Xiao Liu , Xiaosheng Yan , Zhishan Zhou , Rui Huang

Food computing is both important and challenging in computer vision (CV). It significantly contributes to the development of CV algorithms due to its frequent presence in datasets across various applications, ranging from classification and…

We introduce HOT3D, a publicly available dataset for egocentric hand and object tracking in 3D. The dataset offers over 833 minutes (more than 3.7M images) of multi-view RGB/monochrome image streams showing 19 subjects interacting with 33…

Human body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, which are subtle, low-intensity body movements with promising…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Kun Li , Pengyu Liu , Dan Guo , Fei Wang , Zhiliang Wu , Hehe Fan , Meng Wang

With the emergence of LLMs and their integration with other data modalities, multi-modal 3D perception attracts more attention due to its connectivity to the physical world and makes rapid progress. However, limited by existing datasets,…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Ruiyuan Lyu , Jingli Lin , Tai Wang , Shuai Yang , Xiaohan Mao , Yilun Chen , Runsen Xu , Haifeng Huang , Chenming Zhu , Dahua Lin , Jiangmiao Pang

Part-level Action Parsing aims at part state parsing for boosting action recognition in videos. Despite of dramatic progresses in the area of video classification research, a severe problem faced by the community is that the detailed…

计算机视觉与模式识别 · 计算机科学 2021-11-08 Xuanhan Wang , Xiaojia Chen , Lianli Gao , Lechao Chen , Jingkuan Song

We introduce Generalizable 3D-Language Feature Fields (g3D-LF), a 3D representation model pre-trained on large-scale 3D-language dataset for embodied tasks. Our g3D-LF processes posed RGB-D images from agents to encode feature fields for:…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Zihan Wang , Gim Hee Lee

Gait recognition has a rapid development in recent years. However, gait recognition in the wild is not well explored yet. An obvious reason could be ascribed to the lack of diverse training data from the perspective of intrinsic and…

计算机视觉与模式识别 · 计算机科学 2021-06-02 Pengyi Zhang , Huanzhang Dou , Wenhu Zhang , Yuhan Zhao , Songyuan Li , Zequn Qin , Xi Li

The task of text2motion is to generate human motion sequences from given textual descriptions, where the model explores diverse mappings from natural language instructions to human body movements. While most existing works are confined to…

人工智能 · 计算机科学 2024-03-27 Kunhang Li , Yansong Feng

Recognizing human actions in untrimmed videos is an important challenging task. An effective 3D motion representation and a powerful learning model are two key factors influencing recognition performance. In this paper we introduce a new…

计算机视觉与模式识别 · 计算机科学 2018-12-31 Huy-Hieu Pham , Louahdi Khoudour , Alain Crouzil , Pablo Zegers , Sergio A. Velastin

Fine-grained understanding of human actions and poses in videos is essential for human-centric AI applications. In this work, we introduce ActionArt, a fine-grained video-caption dataset designed to advance research in human-centric…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Yi-Xing Peng , Qize Yang , Yu-Ming Tang , Shenghao Fu , Kun-Yu Lin , Xihan Wei , Wei-Shi Zheng

Fatigue detection is valued for people to keep mental health and prevent safety accidents. However, detecting facial fatigue, especially mild fatigue in the real world via machine vision is still a challenging issue due to lack of non-lab…

计算机视觉与模式识别 · 计算机科学 2021-04-22 Zeyu Chen , Xinhang Zhang , Juan Li , Jingxuan Ni , Gang Chen , Shaohua Wang , Fangfang Fan , Changfeng Charles Wang , Xiaotao Li

In human activity recognition (HAR), activity labels have typically been encoded in one-hot format, which has a recent shift towards using textual representations to provide contextual knowledge. Here, we argue that HAR should be anchored…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Shuheng Li , Jiayun Zhang , Xiaohan Fu , Xiyuan Zhang , Jingbo Shang , Rajesh K. Gupta

3D affordance reasoning is essential in associating human instructions with the functional regions of 3D objects, facilitating precise, task-oriented manipulations in embodied AI. However, current methods, which predominantly depend on…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Zeming Wei , Junyi Lin , Yang Liu , Weixing Chen , Jingzhou Luo , Guanbin Li , Liang Lin

Markerless motion capture is an active research in 3D virtualization. In proposed work we presented a system for markerless motion capture for 3D human character animation, paper presents a survey on motion and skeleton tracking techniques…

图形学 · 计算机科学 2014-02-12 Ashish Shingade , Archana Ghotkar

Unconditional human image generation is an important task in vision and graphics, which enables various applications in the creative industry. Existing studies in this field mainly focus on "network engineering" such as designing new…

计算机视觉与模式识别 · 计算机科学 2022-04-26 Jianglin Fu , Shikai Li , Yuming Jiang , Kwan-Yee Lin , Chen Qian , Chen Change Loy , Wayne Wu , Ziwei Liu
‹ 上一页 1 8 9 10 下一页 ›