中文
相关论文

相关论文: Beyond Appearance: Transformer-based Person Identi…

200 篇论文

Context plays a significant role in the generation of motion for dynamic agents in interactive environments. This work proposes a modular method that utilises a learned model of the environment for motion prediction. This modularity…

机器学习 · 计算机科学 2021-01-05 Todor Davchev , Michael Burke , Subramanian Ramamoorthy

Humans appear to represent objects for intuitive physics with coarse, volumetric bodies'' that smooth concavities - trading fine visual details for efficient physical predictions - yet their internal structure is largely unknown.…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Andrey Gizdov , Andrea Procopio , Yichen Li , Daniel Harari , Tomer Ullman

Predicting the future movements of surrounding vehicles is essential for ensuring the safe operation and efficient navigation of autonomous vehicles (AVs) in urban traffic environments. Existing vehicle trajectory prediction methods…

机器人学 · 计算机科学 2025-12-10 Yuansheng Lian , Ke Zhang , Meng Li

Artificial Intelligence (AI) has demonstrated unprecedented performance across various domains, and its application to communication systems is an active area of research. While current methods focus on task-specific solutions, the broader…

This study investigates multimodal turn-taking prediction within human-agent interactions (HAI), particularly focusing on cooperative gaming environments. It comprises both model development and subsequent user study, aiming to refine our…

人机交互 · 计算机科学 2025-03-24 Young-Ho Bae , Casey C. Bennett

Movement synchrony reflects the coordination of body movements between interacting dyads. The estimation of movement synchrony has been automated by powerful deep learning models such as transformer networks. However, instead of designing a…

计算机视觉与模式识别 · 计算机科学 2022-08-03 Jicheng Li , Anjana Bhat , Roghayeh Barmaki

Conventional approaches to human mesh recovery predominantly employ a region-based strategy. This involves initially cropping out a human-centered region as a preprocessing step, with subsequent modeling focused on this zoomed-in image.…

计算机视觉与模式识别 · 计算机科学 2024-02-27 Zeyu Wang , Zhenzhen Weng , Serena Yeung-Levy

Aging presents a significant challenge in face recognition, as changes in skin texture and tone can alter facial features over time, making it particularly difficult to compare images of the same individual taken years apart, such as in…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Pritesh Prakash , Anoop Kumar Rai

As a fundamental aspect of human life, two-person interactions contain meaningful information about people's activities, relationships, and social settings. Human action recognition serves as the foundation for many smart applications, with…

计算机视觉与模式识别 · 计算机科学 2024-05-15 Yao Liu , Gangfeng Cui , Jiahui Luo , Xiaojun Chang , Lina Yao

Automated personality and soft skill assessment from multimodal behavioral data remains challenging due to limited datasets and methods that fail to capture geometric structure inherent in human traits. We introduce RecruitView, a dataset…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Amit Kumar Gupta , Farhan Sheth , Hammad Shaikh , Dheeraj Kumar , Angkul Puniya , Deepak Panwar , Sandeep Chaurasia , Priya Mathur

Accurately modeling affect dynamics, which refers to the changes and fluctuations in emotions and affective displays during human conversations, is crucial for understanding human interactions. By analyzing affect dynamics, we can gain…

计算机视觉与模式识别 · 计算机科学 2023-05-23 Yubin Kim , Dong Won Lee , Paul Pu Liang , Sharifa Algohwinem , Cynthia Breazeal , Hae Won Park

Transformer with self-attention has achieved great success in the area of nature language processing. Recently, there have been a few studies on transformer for end-to-end speech recognition, while its application for hybrid acoustic model…

音频与语音处理 · 电气工程与系统科学 2019-10-24 Liang Lu

Achieving high-performance in multi-object tracking algorithms heavily relies on modeling spatio-temporal relationships during the data association stage. Mainstream approaches encompass rule-based and deep learning-based methods for…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Zhonglin Liu , Shujie Chen , Jianfeng Dong , Xun Wang , Di Zhou

Turn-taking management is crucial for any social interaction. Still, it is challenging to model human-machine interaction due to the complexity of the social context and its multimodal nature. Unlike conventional systems based on silence…

计算与语言 · 计算机科学 2025-06-05 Takeshi Saga , Catherine Pelachaud

Predicting accurate future trajectories of multiple agents is essential for autonomous systems, but is challenging due to the complex agent interaction and the uncertainty in each agent's future behavior. Forecasting multi-agent…

人工智能 · 计算机科学 2021-10-08 Ye Yuan , Xinshuo Weng , Yanglan Ou , Kris Kitani

Emotion recognition is a topic of significant interest in assistive robotics due to the need to equip robots with the ability to comprehend human behavior, facilitating their effective interaction in our society. Consequently, efficient and…

We propose a novel framework for multi-person 3D motion trajectory prediction. Our key observation is that a human's action and behaviors may highly depend on the other persons around. Thus, instead of predicting each human pose trajectory…

计算机视觉与模式识别 · 计算机科学 2021-11-24 Jiashun Wang , Huazhe Xu , Medhini Narasimhan , Xiaolong Wang

It is a challenging task to learn rich and multi-scale spatiotemporal semantics from high-dimensional videos, due to large local redundancy and complex global dependency between video frames. The recent advances in this research have been…

计算机视觉与模式识别 · 计算机科学 2022-02-09 Kunchang Li , Yali Wang , Peng Gao , Guanglu Song , Yu Liu , Hongsheng Li , Yu Qiao

Amodal object completion is a complex task that involves predicting the invisible parts of an object based on visible segments and background information. Learning shape priors is crucial for effective amodal completion, but traditional…

计算机视觉与模式识别 · 计算机科学 2024-05-31 Jianxiong Gao , Xuelin Qian , Longfei Liang , Junwei Han , Yanwei Fu

Classical person re-identification approaches assume that a person of interest has appeared across different cameras and can be queried by one of the existing images. However, in real-world surveillance scenarios, frequently no visual…

计算机视觉与模式识别 · 计算机科学 2020-03-03 Ammarah Farooq , Muhammad Awais , Fei Yan , Josef Kittler , Ali Akbari , Syed Safwan Khalid