中文
相关论文

相关论文: The 2026 ACII Dyadic Conversations (DaiKon) Worksh…

200 篇论文

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

计算与语言 · 计算机科学 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota

Human Object Interaction (HOI) detection aims to localize and infer the relationships between a human and an object. Arguably, training supervised models for this task from scratch presents challenges due to the performance drop over rare…

计算机视觉与模式识别 · 计算机科学 2023-09-08 Ting Lei , Fabian Caba , Qingchao Chen , Hailin Jin , Yuxin Peng , Yang Liu

This paper documents and theorises a self-reinforcing dynamic between two measurable trends: the exponential expansion of large language model (LLM) context windows and the secular contraction of human sustained-attention capacity. We term…

计算与语言 · 计算机科学 2026-03-31 Netanel Eliav

Psychomotor retardation in depression has been associated with speech timing changes from dyadic clinical interviews. In this work, we investigate speech timing features from free-living dyadic interactions. Apart from the possibility of…

声音 · 计算机科学 2022-09-09 Bishal Lamichhane , Nidal Moukaddam , Ankit B. Patel , Ashutosh Sabharwal

Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However, most prior works on multimodal social cues focus predominantly on single-person behaviors or…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Sangmin Lee , Bolin Lai , Fiona Ryan , Bikram Boote , James M. Rehg

Conversational human-AI interaction (CHAI) have recently driven mainstream adoption of AI. However, CHAI poses two key challenges for designers and researchers: users frequently have ambiguous goals and an incomplete understanding of AI…

人机交互 · 计算机科学 2025-01-31 Arthur Caetano , Kavya Verma , Atieh Taheri , Radha Kumaran , Zichen Chen , Jiaao Chen , Tobias Höllerer , Misha Sra

Classifying the general intent of the user utterance in a conversation, also known as Dialogue Act (DA), e.g., open-ended question, statement of opinion, or request for an opinion, is a key step in Natural Language Understanding (NLU) for…

计算与语言 · 计算机科学 2020-05-29 Ali Ahmadvand , Jason Ingyu Choi , Eugene Agichtein

We present a framework for modeling interactional communication in dyadic conversations: given multimodal inputs of a speaker, we autoregressively output multiple possibilities of corresponding listener motion. We combine the motion and…

计算机视觉与模式识别 · 计算机科学 2022-04-19 Evonne Ng , Hanbyul Joo , Liwen Hu , Hao Li , Trevor Darrell , Angjoo Kanazawa , Shiry Ginosar

This paper introduces UDIVA, a new non-acted dataset of face-to-face dyadic interactions, where interlocutors perform competitive and collaborative tasks with different behavior elicitation and cognitive workload. The dataset consists of…

Conversational Speech Synthesis (CSS) aims to effectively take the multimodal dialogue history (MDH) to generate speech with appropriate conversational prosody for target utterance. The key challenge of CSS is to model the interaction…

计算与语言 · 计算机科学 2024-12-30 Zhenqi Jia , Rui Liu

Drones have been widely used in many areas of our daily lives. It relieves people of the burden of holding a controller all the time and makes drone control easier to use for people with disabilities or occupied hands. However, the control…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Xinyi Wang , Xuan Cui , Danxu Li , Fang Liu , Licheng Jiao

Continuous affect prediction involves the discrete time-continuous regression of affect dimensions. Dimensions to be predicted often include arousal and valence. Continuous affect prediction researchers are now embracing multimodal model…

人机交互 · 计算机科学 2020-01-24 Jonny O'Dwyer

Human emotional expression emerges through coordinated vocal, facial, and gestural signals. While speech face alignment is well established, the broader dynamics linking emotionally expressive speech to regional facial and hand motion…

多媒体 · 计算机科学 2025-06-13 Von Ralph Dane Marquez Herbuela , Yukie Nagai

Surveys and interviews are widely used for collecting insights on emerging or hypothetical scenarios. Traditional human-led methods often face challenges related to cost, scalability, and consistency. Recently, various domains have begun to…

人机交互 · 计算机科学 2025-03-05 Jiangbo Yu , Jinhua Zhao , Luis Miranda-Moreno , Matthew Korp

Automatic coding patient behaviors is essential to support decision making for psychotherapists during the motivational interviewing (MI), a collaborative communication intervention approach to address psychiatric issues, such as alcohol…

计算与语言 · 计算机科学 2024-03-26 Guangzeng Han , Weisi Liu , Xiaolei Huang , Brian Borsari

Estimating the treatment effect within network structures is a key focus in online controlled experiments, particularly for social media platforms. We investigate a scenario where the unit-level outcome of interest comprises a series of…

统计方法学 · 统计学 2025-05-28 Yilin Li , Lu Deng , Yong Wang , Wang Miao

Content moderation is a central mechanism through which platforms attempt to balance user engagement with community governance. Yet existing research has largely treated moderation as a uniform intervention, overlooking how moderator…

计算机与社会 · 计算机科学 2026-05-18 Siyi Zhou , Lindsay Young , Marlon Twyman , Emilio Ferrara

There is growing concern that AI chatbots might fuel delusional beliefs in users. Some have suggested that humans and chatbots mutually reinforce false beliefs over time, but quantitative evidence is lacking. Using a unique dataset of chat…

计算与语言 · 计算机科学 2026-04-29 Ashish Mehta , Jared Moore , Jacy Reese Anthis , William Agnew , Eric Lin , Peggy Yin , Desmond C. Ong , Nick Haber , Carol Dweck

This paper addresses the gap in predicting turn-taking and backchannel actions in human-machine conversations using multi-modal signals (linguistic, acoustic, and visual). To overcome the limitation of existing datasets, we propose an…

计算与语言 · 计算机科学 2025-05-21 Yuxin Lin , Yinglin Zheng , Ming Zeng , Wangzheng Shi

Multimodal dialogue emotion recognition captures emotional cues by fusing text, visual, and audio modalities. However, existing approaches still suffer from notable limitations in modeling emotional dependencies and learning multimodal…

多媒体 · 计算机科学 2026-03-12 Yunsheng Wang , Yuntao Shou , Yilong Tan , Wei Ai , Tao Meng , Keqin Li