English
Related papers

Related papers: The 2026 ACII Dyadic Conversations (DaiKon) Worksh…

200 papers

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

Computation and Language · Computer Science 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota

Human Object Interaction (HOI) detection aims to localize and infer the relationships between a human and an object. Arguably, training supervised models for this task from scratch presents challenges due to the performance drop over rare…

Computer Vision and Pattern Recognition · Computer Science 2023-09-08 Ting Lei , Fabian Caba , Qingchao Chen , Hailin Jin , Yuxin Peng , Yang Liu

This paper documents and theorises a self-reinforcing dynamic between two measurable trends: the exponential expansion of large language model (LLM) context windows and the secular contraction of human sustained-attention capacity. We term…

Computation and Language · Computer Science 2026-03-31 Netanel Eliav

Psychomotor retardation in depression has been associated with speech timing changes from dyadic clinical interviews. In this work, we investigate speech timing features from free-living dyadic interactions. Apart from the possibility of…

Sound · Computer Science 2022-09-09 Bishal Lamichhane , Nidal Moukaddam , Ankit B. Patel , Ashutosh Sabharwal

Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However, most prior works on multimodal social cues focus predominantly on single-person behaviors or…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Sangmin Lee , Bolin Lai , Fiona Ryan , Bikram Boote , James M. Rehg

Conversational human-AI interaction (CHAI) have recently driven mainstream adoption of AI. However, CHAI poses two key challenges for designers and researchers: users frequently have ambiguous goals and an incomplete understanding of AI…

Human-Computer Interaction · Computer Science 2025-01-31 Arthur Caetano , Kavya Verma , Atieh Taheri , Radha Kumaran , Zichen Chen , Jiaao Chen , Tobias Höllerer , Misha Sra

Classifying the general intent of the user utterance in a conversation, also known as Dialogue Act (DA), e.g., open-ended question, statement of opinion, or request for an opinion, is a key step in Natural Language Understanding (NLU) for…

Computation and Language · Computer Science 2020-05-29 Ali Ahmadvand , Jason Ingyu Choi , Eugene Agichtein

We present a framework for modeling interactional communication in dyadic conversations: given multimodal inputs of a speaker, we autoregressively output multiple possibilities of corresponding listener motion. We combine the motion and…

Computer Vision and Pattern Recognition · Computer Science 2022-04-19 Evonne Ng , Hanbyul Joo , Liwen Hu , Hao Li , Trevor Darrell , Angjoo Kanazawa , Shiry Ginosar

This paper introduces UDIVA, a new non-acted dataset of face-to-face dyadic interactions, where interlocutors perform competitive and collaborative tasks with different behavior elicitation and cognitive workload. The dataset consists of…

Conversational Speech Synthesis (CSS) aims to effectively take the multimodal dialogue history (MDH) to generate speech with appropriate conversational prosody for target utterance. The key challenge of CSS is to model the interaction…

Computation and Language · Computer Science 2024-12-30 Zhenqi Jia , Rui Liu

Drones have been widely used in many areas of our daily lives. It relieves people of the burden of holding a controller all the time and makes drone control easier to use for people with disabilities or occupied hands. However, the control…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Xinyi Wang , Xuan Cui , Danxu Li , Fang Liu , Licheng Jiao

Continuous affect prediction involves the discrete time-continuous regression of affect dimensions. Dimensions to be predicted often include arousal and valence. Continuous affect prediction researchers are now embracing multimodal model…

Human-Computer Interaction · Computer Science 2020-01-24 Jonny O'Dwyer

Human emotional expression emerges through coordinated vocal, facial, and gestural signals. While speech face alignment is well established, the broader dynamics linking emotionally expressive speech to regional facial and hand motion…

Multimedia · Computer Science 2025-06-13 Von Ralph Dane Marquez Herbuela , Yukie Nagai

Surveys and interviews are widely used for collecting insights on emerging or hypothetical scenarios. Traditional human-led methods often face challenges related to cost, scalability, and consistency. Recently, various domains have begun to…

Human-Computer Interaction · Computer Science 2025-03-05 Jiangbo Yu , Jinhua Zhao , Luis Miranda-Moreno , Matthew Korp

Automatic coding patient behaviors is essential to support decision making for psychotherapists during the motivational interviewing (MI), a collaborative communication intervention approach to address psychiatric issues, such as alcohol…

Computation and Language · Computer Science 2024-03-26 Guangzeng Han , Weisi Liu , Xiaolei Huang , Brian Borsari

Estimating the treatment effect within network structures is a key focus in online controlled experiments, particularly for social media platforms. We investigate a scenario where the unit-level outcome of interest comprises a series of…

Methodology · Statistics 2025-05-28 Yilin Li , Lu Deng , Yong Wang , Wang Miao

Content moderation is a central mechanism through which platforms attempt to balance user engagement with community governance. Yet existing research has largely treated moderation as a uniform intervention, overlooking how moderator…

Computers and Society · Computer Science 2026-05-18 Siyi Zhou , Lindsay Young , Marlon Twyman , Emilio Ferrara

There is growing concern that AI chatbots might fuel delusional beliefs in users. Some have suggested that humans and chatbots mutually reinforce false beliefs over time, but quantitative evidence is lacking. Using a unique dataset of chat…

Computation and Language · Computer Science 2026-04-29 Ashish Mehta , Jared Moore , Jacy Reese Anthis , William Agnew , Eric Lin , Peggy Yin , Desmond C. Ong , Nick Haber , Carol Dweck

This paper addresses the gap in predicting turn-taking and backchannel actions in human-machine conversations using multi-modal signals (linguistic, acoustic, and visual). To overcome the limitation of existing datasets, we propose an…

Computation and Language · Computer Science 2025-05-21 Yuxin Lin , Yinglin Zheng , Ming Zeng , Wangzheng Shi

Multimodal dialogue emotion recognition captures emotional cues by fusing text, visual, and audio modalities. However, existing approaches still suffer from notable limitations in modeling emotional dependencies and learning multimodal…

Multimedia · Computer Science 2026-03-12 Yunsheng Wang , Yuntao Shou , Yilong Tan , Wei Ai , Tao Meng , Keqin Li