中文
相关论文

相关论文: SocialGesture: Delving into Multi-person Gesture U…

200 篇论文

Social group detection, or the identification of humans involved in reciprocal interpersonal interactions (e.g., family members, friends, and customers and merchants), is a crucial component of social intelligence needed for agents…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Jeffri Murrugarra-Llerena , Pranav Chitale , Zicheng Liu , Kai Ao , Yujin Ham , Guha Balakrishnan , Paola Cascante-Bonilla

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

Autism diagnosis presents a major challenge due to the vast heterogeneity of the condition and the elusive nature of early detection. Atypical gait and gesture patterns are dominant behavioral characteristics of autism and can provide…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Sania Zahan , Zulqarnain Gilani , Ghulam Mubashar Hassan , Ajmal Mian

Large language models (LLMs) have rapidly evolved from general-purpose systems to multimodal models capable of processing text, images, and audio. As both general-purpose LLMs (GLLMs) and multimodal LLMs (MLLMs) gain widespread adoption,…

软件工程 · 计算机科学 2026-04-08 Yujian Liu , Xiao Yu , Jacky Keung , Xing Hu , Xin Xia , Xiaoxue Ma

Multimodal large language models (MLLMs) have shown remarkable performance in vision-language tasks. However, existing MLLMs are primarily trained on generic datasets, limiting their ability to reason on domain-specific visual cues such as…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Hatef Otroshi Shahreza , Sébastien Marcel

Co-speech gestures play a vital role in non-verbal communication. In this paper, we introduce a new framework for co-speech gesture understanding in the wild. Specifically, we propose three new tasks and benchmarks to evaluate a model's…

计算机视觉与模式识别 · 计算机科学 2025-08-22 Sindhu B Hegde , K R Prajwal , Taein Kwon , Andrew Zisserman

In the field of affective computing, researchers in the community have promoted the performance of models and algorithms by using the complementarity of multimodal information. However, the emergence of more and more modal information makes…

计算机视觉与模式识别 · 计算机科学 2023-05-01 Binqiang Wang , Gang Dong , Yaqian Zhao , Rengang Li , Lu Cao , Lihua Lu

Multi-modal large language models have garnered significant interest recently. Though, most of the works focus on vision-language multi-modal models providing strong capabilities in following vision-and-language instructions. However, we…

计算与语言 · 计算机科学 2023-09-19 Yu Shu , Siwei Dong , Guangyao Chen , Wenhao Huang , Ruihua Zhang , Daochen Shi , Qiqi Xiang , Yemin Shi

Human-to-Robot handovers are useful for many Human-Robot Interaction scenarios. It is important to recognize when a human intends to initiate handovers, so that the robot does not try to take objects from humans when a handover is not…

计算机视觉与模式识别 · 计算机科学 2021-01-01 Jun Kwan , Chinkye Tan , Akansel Cosgun

As AI becomes more closely integrated with peoples' daily activities, socially intelligent AI that can understand and interact seamlessly with humans in daily lives is increasingly important. However, current works in AI social reasoning…

计算与语言 · 计算机科学 2025-12-02 Hengzhi Li , Megan Tjandrasuwita , Yi R. Fung , Armando Solar-Lezama , Paul Pu Liang

Image-text retrieval of natural scenes has been a popular research topic. Since image and text are heterogeneous cross-modal data, one of the key challenges is how to learn comprehensive yet unified representations to express the…

计算机视觉与模式识别 · 计算机科学 2019-10-14 Sijin Wang , Ruiping Wang , Ziwei Yao , Shiguang Shan , Xilin Chen

Existing gesture interfaces only work with a fixed set of gestures defined either by interface designers or by users themselves, which introduces learning or demonstration efforts that diminish their naturalness. Humans, on the other hand,…

计算与语言 · 计算机科学 2024-11-05 Xin Zeng , Xiaoyu Wang , Tengxiang Zhang , Chun Yu , Shengdong Zhao , Yiqiang Chen

Gesture synthesis has gained significant attention as a critical research field, aiming to produce contextually appropriate and natural gestures corresponding to speech or textual input. Although deep learning-based approaches have achieved…

计算与语言 · 计算机科学 2024-05-29 Nan Gao , Zeyu Zhao , Zhi Zeng , Shuwu Zhang , Dongdong Weng , Yihua Bao

The proliferation of social network data has unlocked unprecedented opportunities for extensive, data-driven exploration of human behavior. The structural intricacies of social networks offer insights into various computational social…

社会与信息网络 · 计算机科学 2024-01-03 Julie Jiang , Emilio Ferrara

When humans speak, gestures help convey communicative intentions, such as adding emphasis or describing concepts. However, current co-speech gesture generation methods rely solely on superficial linguistic cues (e.g. speech audio or text…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Pinxin Liu , Haiyang Liu , Luchuan Song , Jason J. Corso , Chenliang Xu

Human action recognition as an important application of computer vision has been studied for decades. Among various approaches, skeleton-based methods recently attract increasing attention due to their robust and superior performance.…

计算机视觉与模式识别 · 计算机科学 2021-02-26 Tingtian Li , Zixun Sun , Xiao Chen

Humans have long been recorded in a variety of forms since antiquity. For example, sculptures and paintings were the primary media for depicting human beings before the invention of cameras. However, most current human-centric computer…

计算机视觉与模式识别 · 计算机科学 2023-04-06 Xuan Ju , Ailing Zeng , Jianan Wang , Qiang Xu , Lei Zhang

Hand gestures have evolved into a natural and intuitive means of engaging with technology. The objective of this research is to develop a robust system that can accurately recognize and classify hand gestures representing numbers. The…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Sangeetha K , Balaji VS , Kamalesh P , Anirudh Ganapathy PS

Memes are a dominant medium for online communication and manipulation because meaning emerges from interactions between embedded text, imagery, and cultural context. Existing meme research is distributed across tasks (hate, misogyny,…

The development of Large Vision-Language Models (LVLMs) is striving to catch up with the success of Large Language Models (LLMs), yet it faces more challenges to be resolved. Very recent works enable LVLMs to localize object-level visual…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Zhipeng Huang , Zhizheng Zhang , Zheng-Jun Zha , Yan Lu , Baining Guo