English
Related papers

Related papers: The KIT Motion-Language Dataset

200 papers

Understanding human behavior is key for robots and intelligent systems that share a space with people. Accordingly, research that enables such systems to perceive, track, learn and predict human behavior as well as to plan and interact with…

The annotation of image and video data of large datasets is a fundamental task in multimedia information retrieval and computer vision applications. In order to support the users during the image and video annotation process, several…

Computer Vision and Pattern Recognition · Computer Science 2015-02-19 Gianluigi Ciocca , Paolo Napoletano , Raimondo Schettini

Despite growing interest in using large language models (LLMs) to automate annotation, their effectiveness in complex, nuanced, and multi-dimensional labelling tasks remains relatively underexplored. This study focuses on annotation for the…

Information Retrieval · Computer Science 2025-07-02 Leila Tavakoli , Hamed Zamani

Recent advances in multimodal Human-Robot Interaction (HRI) datasets emphasize the integration of speech and gestures, allowing robots to absorb explicit knowledge and tacit understanding. However, existing datasets primarily focus on…

Robotics · Computer Science 2025-02-25 Snehesh Shrestha , Yantian Zha , Saketh Banagiri , Ge Gao , Yiannis Aloimonos , Cornelia Fermüller

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeVAn (Dense Video Annotation). The dataset contains 8.5K…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Tingkai Liu , Yunzhe Tao , Haogeng Liu , Qihang Fan , Ding Zhou , Huaibo Huang , Ran He , Hongxia Yang

Although the amount of available spoken content is steadily increasing, extracting information and knowledge from speech recordings remains challenging. Beyond enhancing traditional information retrieval methods such as speech search and…

Information Retrieval · Computer Science 2025-07-08 Sirina Håland , Trond Karlsen Strøm , Petra Galuščáková

Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current models remain severely constrained by the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Jiahao Wang , Yufeng Yuan , Rujie Zheng , Youtian Lin , Jian Gao , Lin-Zhuo Chen , Yajie Bao , Yi Zhang , Chang Zeng , Yanxi Zhou , Xiao-Xiao Long , Hao Zhu , Zhaoxiang Zhang , Xun Cao , Yao Yao

Processing human affective behavior is important for developing intelligent agents that interact with humans in complex interaction scenarios. A large number of current approaches that address this problem focus on classifying emotion…

Human-Computer Interaction · Computer Science 2019-09-02 Pablo Barros , Nikhil Churamani , Angelica Lim , Stefan Wermter

In this work, we present Semantic Gesticulator, a novel framework designed to synthesize realistic gestures accompanying speech with strong semantic correspondence. Semantically meaningful gestures are crucial for effective non-verbal…

Graphics · Computer Science 2025-10-23 Zeyi Zhang , Tenglong Ao , Yuyao Zhang , Qingzhe Gao , Chuan Lin , Baoquan Chen , Libin Liu

Recent years have seen an increasing trend in the volume of personal media captured by users, thanks to the advent of smartphones and smart glasses, resulting in large media collections. Despite conversation being an intuitive…

Computation and Language · Computer Science 2022-11-17 Seungwhan Moon , Satwik Kottur , Alborz Geramifard , Babak Damavandi

Existing representations for human motion, such as MotionGPT, often operate as black-box latent vectors with limited interpretability and build on joint positions which can cause ambiguity. Inspired by the hierarchical structure of natural…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yao Zhang , Zhuchenyang Liu , Yu Xiao

Mobile machines typically working in a closed site, have a high potential to utilize autonomous driving technology. However, vigorously thriving development and innovation are happening mostly in the area of passenger cars. In contrast,…

Computer Vision and Pattern Recognition · Computer Science 2020-07-09 Yusheng Xiang , Hongzhe Wang , Tianqing Su , Ruoyu Li , Christine Brach , Samuel S. Mao , Marcus Geimer

We introduce the concept of notational animating, an interaction paradigm for animation authoring where users sketch high-level notations over static drawings to indicate intended motions, which are then interpreted by automatic methods…

Human-Computer Interaction · Computer Science 2026-03-10 Xinyu Shi , Li-Yi Wei , Nanxuan Zhao , Jian Zhao , Rubaiat Habib Kazi

We report on novel investigations into training models that make sentences concise. We define the task and show that it is different from related tasks such as summarization and simplification. For evaluation, we release two test sets,…

Computation and Language · Computer Science 2022-11-09 Felix Stahlberg , Aashish Kumar , Chris Alberti , Shankar Kumar

Intelligent instruction-following robots capable of improving from autonomously collected experience have the potential to transform robot learning: instead of collecting costly teleoperated demonstration data, large-scale deployment of…

Robotics · Computer Science 2025-02-26 Zhiyuan Zhou , Pranav Atreya , Abraham Lee , Homer Walke , Oier Mees , Sergey Levine

We propose an approach for annotating object classes using free-form text written by undirected and untrained annotators. Free-form labeling is natural for annotators, they intuitively provide very specific and exhaustive labels, and no…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Jordi Pont-Tuset , Michael Gygli , Vittorio Ferrari

Human motion generation aims to generate natural human pose sequences and shows immense potential for real-world applications. Substantial progress has been made recently in motion data collection technologies and generation methods, laying…

Computer Vision and Pattern Recognition · Computer Science 2023-11-16 Wentao Zhu , Xiaoxuan Ma , Dongwoo Ro , Hai Ci , Jinlu Zhang , Jiaxin Shi , Feng Gao , Qi Tian , Yizhou Wang

Behavior-related research areas such as motion prediction/planning, representation/imitation learning, behavior modeling/generation, and algorithm testing, require support from high-quality motion datasets containing interactive driving…

While densely annotated image captions significantly facilitate the learning of robust vision-language alignment, methodologies for systematically optimizing human annotation efforts remain underexplored. We introduce Chain-of-Talkers…

Computation and Language · Computer Science 2025-06-03 Yijun Shen , Delong Chen , Fan Liu , Xingyu Wang , Chuanyi Zhang , Liang Yao , Yuhui Zheng

This paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos. We refer to it as HACS (Human Action Clips and Segments). We leverage both consensus and disagreement among…

Computer Vision and Pattern Recognition · Computer Science 2019-09-05 Hang Zhao , Antonio Torralba , Lorenzo Torresani , Zhicheng Yan