English
Related papers

Related papers: Multipar-T: Multiparty-Transformer for Capturing C…

200 papers

The addressee estimation (understanding to whom somebody is talking) is a fundamental task for human activity recognition in multi-party conversation scenarios. Specifically, in the field of human-robot interaction, it becomes even more…

Artificial Intelligence · Computer Science 2025-02-03 Iveta Bečková , Štefan Pócoš , Giulia Belgiovine , Marco Matarese , Omar Eldardeer , Alessandra Sciutti , Carlo Mazzola

Human activity recognition in videos has been widely studied and has recently gained significant advances with deep learning approaches; however, it remains a challenging task. In this paper, we propose a novel framework that simultaneously…

Computer Vision and Pattern Recognition · Computer Science 2021-01-25 Dong-Gyu Lee , Seong-Whan Lee

Predicting pedestrian crossing intention is crucial for autonomous vehicles to prevent pedestrian-related collisions. However, effectively extracting and integrating complementary cues from different types of data remains one of the major…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yuanzhe Li , Steffen Müller

In this paper, we introduce \texttt{IAFormer}, a novel Transformer-based architecture that efficiently integrates pairwise particle interactions through a dynamic sparse attention mechanism. \texttt{IAFormer} has two new mechanisms within…

High Energy Physics - Phenomenology · Physics 2026-04-21 W. Esmail , A. Hammad , M. Nojiri

Multimodal sentiment analysis is an important research area that predicts speaker's sentiment tendency through features extracted from textual, visual and acoustic modalities. The central challenge is the fusion method of the multimodal…

Computation and Language · Computer Science 2020-09-29 Zilong Wang , Zhaohong Wan , Xiaojun Wan

Some group activities, such as team sports and choreographed dances, involve closely coupled interaction between participants. Here we investigate the tasks of inferring and predicting participant behavior, in terms of motion paths and…

Computer Vision and Pattern Recognition · Computer Science 2022-07-15 Bo Hu , Tat-Jen Cham

Conversational engagement estimation is posed as a regression problem, entailing the identification of the favorable attention and involvement of the participants in the conversation. This task arises as a crucial pursuit to gain insights…

Multimodal emotion recognition (MER) aims to infer human affect by jointly modeling audio and visual cues; however, existing approaches often struggle with temporal misalignment, weakly discriminative feature representations, and suboptimal…

Multimedia · Computer Science 2026-01-21 Joe Dhanith P R , Shravan Venkatraman , Vigya Sharma , Santhosh Malarvannan

This report outlines the use of a relational representation in a Multi-Agent domain to model the behaviour of the whole system. A desired property in this systems is the ability of the team members to work together to achieve a common goal…

Artificial Intelligence · Computer Science 2010-11-01 Grazia Bombini , Raquel Ros , Stefano Ferilli , Ramon Lopez de Mantaras

Existing voice AI assistants treat every detected pause as an invitation to speak. This works in dyadic dialogue, but in multi-party settings, where an AI assistant participates alongside multiple speakers, pauses are abundant and…

Artificial Intelligence · Computer Science 2026-03-13 Kratika Bhagtani , Mrinal Anand , Yu Chen Xu , Amit Kumar Singh Yadav

Human intention prediction is a growing area of research where an activity in a video has to be anticipated by a vision-based system. To this end, the model creates a representation of the past, and subsequently, it produces future…

Computer Vision and Pattern Recognition · Computer Science 2022-10-27 Nada Osman , Guglielmo Camporese , Lamberto Ballan

We tackle a challenging task: multi-view and multi-modal event detection that detects events in a wide-range real environment by utilizing data from distributed cameras and microphones and their weak labels. In this task, distributed…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-21 Masahiro Yasuda , Yasunori Ohishi , Shoichiro Saito , Noboru Harada

Aggregating multi-modality data to obtain reliable data representation attracts more and more attention. Recent studies demonstrate that Transformer models usually work well for multi-modality tasks. Existing Transformers generally either…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Xixi Wang , Xiao Wang , Bo Jiang , Jin Tang , Bin Luo

Multi-modal multi-view action recognition is a rapidly growing field in computer vision, offering significant potential for applications in surveillance. However, current datasets often fail to address real-world challenges such as…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Trung Thanh Nguyen , Yasutomo Kawanishi , Vijay John , Takahiro Komamizu , Ichiro Ide

Multi-agent trajectory prediction is a fundamental problem in autonomous driving. The key challenges in prediction are accurately anticipating the behavior of surrounding agents and understanding the scene context. To address these…

Computer Vision and Pattern Recognition · Computer Science 2022-03-04 Elmira Amirloo , Amir Rasouli , Peter Lakner , Mohsen Rohani , Jun Luo

This study examines how Critical Care Air Transport Team (CCATT) members are trained using mixed-reality simulations that replicate the high-pressure conditions of aeromedical evacuation. Each team - a physician, nurse, and respiratory…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Divya Mereddy , Marcos Quinones-Grueiro , Ashwin T S , Eduardo Davalos , Gautam Biswas , Kent Etherton , Tyler Davis , Katelyn Kay , Jill Lear , Benjamin Goldberg

In Click-through rate (CTR) prediction models, a user's interest is usually represented as a fixed-length vector based on her history behaviors. Recently, several methods are proposed to learn an attentive weight for each user behavior and…

Information Retrieval · Computer Science 2022-10-28 Zuowu Zheng , Xiaofeng Gao , Junwei Pan , Qi Luo , Guihai Chen , Dapeng Liu , Jie Jiang

Humans express their emotions via facial expressions, voice intonation and word choices. To infer the nature of the underlying emotion, recognition models may use a single modality, such as vision, audio, and text, or a combination of…

Machine Learning · Computer Science 2022-02-21 Vandana Rajan , Alessio Brutti , Andrea Cavallaro

Multiparty session types are designed to abstractly capture the structure of communication protocols and verify behavioural properties. One important such property is progress, i.e., the absence of deadlock. Distributed algorithms often…

Logic in Computer Science · Computer Science 2024-02-14 Kirstin Peters , Uwe Nestmann , Christoph Wagner

Multi-person motion prediction is a challenging task, especially for real-world scenarios of highly interacted persons. Most previous works have been devoted to studying the case of weak interactions (e.g., walking together), in which…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Yanwen Fang , Jintai Chen , Peng-Tao Jiang , Chao Li , Yifeng Geng , Eddy K. F. Lam , Guodong Li