English
Related papers

Related papers: BFMD: A Full-Match Badminton Dense Dataset for Den…

200 papers

Linking human motion and natural language is of great interest for the generation of semantic representations of human activities as well as for the generation of robot activities based on natural language input. However, while there have…

Robotics · Computer Science 2018-08-10 Matthias Plappert , Christian Mandery , Tamim Asfour

Characterizing and quantifying gender representation disparities in audiovisual storytelling contents is necessary to grasp how stereotypes may perpetuate on screen. In this article, we consider the high-level construct of objectification…

Large Multimodal Models (LMMs) have demonstrated exceptional performance in video captioning tasks, particularly for short videos. However, as the length of the video increases, generating long, detailed captions becomes a significant…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Hongchen Wei , Zhihong Tan , Yaosi Hu , Chang Wen Chen , Zhenzhong Chen

Rapid information access is vital during wildfires, yet traditional data sources are slow and costly. Social media offers real-time updates, but extracting relevant insights remains a challenge. In this work, we focus on multimodal wildfire…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Braeden Sherritt , Isar Nejadgholi , Efstratios Aivaliotis , Khaled Mslmani , Marzieh Amini

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able…

In this work, we present a database of multimodal communication features extracted from debate speeches in the 2019 North American Universities Debate Championships (NAUDC). Feature sets were extracted from the visual (facial expression,…

Generative models for audio-conditioned dance motion synthesis map music features to dance movements. Models are trained to associate motion patterns to audio patterns, usually without an explicit knowledge of the human body. This approach…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Davide Moltisanti , Jinyi Wu , Bo Dai , Chen Change Loy

Multi-modal multi-party conversation (MMC) is a less studied yet important topic of research due to that it well fits real-world scenarios and thus potentially has more widely-used applications. Compared with the traditional multi-modal…

Computation and Language · Computer Science 2024-12-24 Yueqian Wang , Xiaojun Meng , Yuxuan Wang , Jianxin Liang , Qun Liu , Dongyan Zhao

Automatic Emotion Detection (ED) aims to build systems to identify users' emotions automatically. This field has the potential to enhance HCI, creating an individualised experience for the user. However, ED systems tend to perform poorly on…

Human-Computer Interaction · Computer Science 2023-07-27 Annanda Sousa , Karen Young , Mathieu D'aquin , Manel Zarrouk , Jennifer Holloway

The dynamic nature of esports makes the situation relatively complicated for average viewers. Esports broadcasting involves game expert casters, but the caster-dependent game commentary is not enough to fully understand the game situation.…

Computation and Language · Computer Science 2024-05-01 Zhihao Zhang , Feiqi Cao , Yingbin Mo , Yiran Zhang , Josiah Poon , Caren Han

Annotating a large-scale in-the-wild person re-identification dataset especially of marathon runners is a challenging task. The variations in the scenarios such as camera viewpoints, resolution, occlusion, and illumination make the problem…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Pranjal Singh Rajput , Yeshwanth Napolean , Jan van Gemert

Multimodal language analysis is a rapidly evolving field that leverages multiple modalities to enhance the understanding of high-level semantics underlying human conversational utterances. Despite its significance, little research has…

Computation and Language · Computer Science 2025-04-25 Hanlei Zhang , Zhuohang Li , Yeshuang Zhu , Hua Xu , Peiwu Wang , Haige Zhu , Jie Zhou , Jinchao Zhang

While existing video benchmarks largely consider specialized downstream tasks like retrieval or question-answering (QA), contemporary multimodal AI systems must be capable of well-rounded common-sense reasoning akin to human visual…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Kate Sanders , Benjamin Van Durme

This study proposes a framework for enhancing the stroke quality of badminton players by generating personalized motion guides, utilizing a multimodal wearable dataset. These guides are based on counterfactual algorithms and aim to reduce…

Human-Computer Interaction · Computer Science 2024-05-21 Minwoo Seong , Gwangbin Kim , Yumin Kang , Junhyuk Jang , Joseph DelPreto , SeungJun Kim

Multi-object tracking (MOT) in team sports is particularly challenging due to the fast-paced motion and frequent occlusions resulting in motion blur and identity switches, respectively. Predicting player positions in such scenarios is…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Dheeraj Khanna , Jerrin Bright , Yuhao Chen , John S. Zelek

We present a novel dataset for training and benchmarking semantic SLAM methods. The dataset consists of 200 long sequences, each one containing 3000-5000 data frames. We generate the sequences using realistic home layouts. For that we…

Computer Vision and Pattern Recognition · Computer Science 2019-09-27 Pavel Kirsanov , Airat Gaskarov , Filipp Konokhov , Konstantin Sofiiuk , Anna Vorontsova , Igor Slinko , Dmitry Zhukov , Sergey Bykov , Olga Barinova , Anton Konushin

Computer vision and video understanding have transformed sports analytics by enabling large-scale, automated analysis of game dynamics from broadcast footage. Despite significant advances in player and ball tracking, pose estimation, action…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Arnau Barrera Roy , Albert Clapés Sintes

Recent advances in multi-modal models have demonstrated strong performance in tasks such as image generation and reasoning. However, applying these models to the fire domain remains challenging due to the lack of publicly available datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Zixuan Liu , Siavash H. Khajavi , Guangkai Jiang

Solving Bengali Math Word Problems (MWPs) remains a major challenge in natural language processing (NLP) due to the language's low-resource status and the multi-step reasoning required. Existing models struggle with complex Bengali MWPs,…

Computation and Language · Computer Science 2025-07-31 Bidyarthi Paul , Jalisha Jashim Era , Mirazur Rahman Zim , Tahmid Sattar Aothoi , Faisal Muhammad Shah

Medical data poses a daunting challenge for AI algorithms: it exists in many different modalities, experiences frequent distribution shifts, and suffers from a scarcity of examples and labels. Recent advances, including transformers and…

‹ Prev 1 8 9 10 Next ›