中文
相关论文

相关论文: SIMMC: Situated Interactive Multi-Modal Conversati…

200 篇论文

With growing capabilities of large language models (LLMs) comes growing affordances for human-like and context-aware conversational partners. On from this, some recent work has investigated the use of LLMs to simulate multiple…

人机交互 · 计算机科学 2023-12-29 Samuel Rhys Cox

A fundamental tension exists between the demand for sophisticated AI assistance in web search and the need for user data privacy. Current centralized models require users to transmit sensitive browsing data to external services, which…

人机交互 · 计算机科学 2026-01-16 Saber Zerhoudi , Michael Granitzer

In online education, innovative tools are crucial for enhancing learning outcomes. SAM (Study with AI Mentor) is an advanced platform that integrates educational videos with a context-aware chat interface powered by large language models.…

人工智能 · 计算机科学 2025-02-25 Anna Bodonhelyi , Enkeleda Thaqi , Süleyman Özdel , Efe Bozkir , Enkelejda Kasneci

We introduce HoME: a Household Multimodal Environment for artificial agents to learn from vision, audio, semantics, physics, and interaction with objects and other agents, all within a realistic context. HoME integrates over 45,000 diverse…

Simulations, although powerful in accurately replicating real-world systems, often remain inaccessible to non-technical users due to their complexity. Conversely, large language models (LLMs) provide intuitive, language-based interactions…

计算与语言 · 计算机科学 2025-05-22 Jacob Kleiman , Kevin Frank , Joseph Voyles , Sindy Campagna

Even in our increasingly text-intensive times, the primary site of language use is situated, co-present interaction. It is primary ontogenetically and phylogenetically, and it is arguably also still primary in negotiating everyday social…

计算与语言 · 计算机科学 2023-02-20 David Schlangen

Multimodal signals, including text, audio, image, and video, can be integrated into Semantic Communication (SC) systems to provide an immersive experience with low latency and high quality at the semantic level. However, the multimodal SC…

人工智能 · 计算机科学 2024-08-06 Feibo Jiang , Li Dong , Yubo Peng , Kezhi Wang , Kun Yang , Cunhua Pan , Xiaohu You

LLM agents are increasingly used for social simulation, yet emotion is often treated as a transient cue, causing emotional amnesia and weak long-horizon continuity. We present Sentipolis, a framework for emotionally stateful agents that…

人工智能 · 计算机科学 2026-04-22 Chiyuan Fu , Lyuhao Chen , Yunze Xiao , Weihao Xuan , Carlos Busso , Mona Diab

In this paper, we explore a multi-task semantic communication (SemCom) system for distributed sources, extending the existing focus on collaborative single-task execution. We build on the cooperative multi-task processing introduced in [1],…

信号处理 · 电气工程与系统科学 2025-06-11 Ahmad Halimi Razlighi , Maximilian H. V. Tillmann , Edgar Beck , Carsten Bockelmann , Armin Dekorsy

Research funding discovery remains fundamentally fragmented: researchers navigate disparate agency portals (e.g., in the United States, NSF, NIH, DARPA, Grants.gov, and many others) with heterogeneous interfaces, search capabilities, and…

人工智能 · 计算机科学 2026-05-05 Zhisheng Tang , Mayank Kejriwal

Engaging in smooth conversations with others is a crucial social skill. However, differences in knowledge between conversation participants can sometimes hinder effective communication. To tackle this issue, this study proposes a real-time…

人机交互 · 计算机科学 2025-06-23 Yuichiro Fujimoto

With the increasing prevalence of multimodal content on social media, sentiment analysis faces significant challenges in effectively processing heterogeneous data and recognizing multi-label emotions. Existing methods often lack effective…

计算与语言 · 计算机科学 2025-08-26 Xilai Xu , Zilin Zhao , Chengye Song , Zining Wang , Jinhe Qiang , Jiongrui Yan , Yuhuai Lin

As Large Language Models (LLMs) transition from static tools to autonomous agents, traditional evaluation benchmarks that measure performance on downstream tasks are becoming insufficient. These methods fail to capture the emergent social…

人工智能 · 计算机科学 2025-10-03 Zarreen Reza

With the upcoming capabilities of integrated sensing and communication (ISAC) and the incorporation of user equipment (UE) like unmanned aerial vehicles (UAVs) in 6G mobile networks, there is a significant opportunity to enhance situational…

Virtual Reality (VR) systems collect fine-grained behavioral and biometric data, yet privacy policies are rarely read or understood due to their complex language, length, and poor integration into users' interaction workflows. To lower the…

人机交互 · 计算机科学 2026-03-03 Vincent Freiberger , Moritz Dresch , Florian Alt , Arthur Fleig , Viktorija Paneva

We devise a multimodal conversation system for dialogue utterances composed of text, image or both modalities. We leverage Auxiliary UnsuperviseD vIsual and TExtual Data (AUDITED). To improve the performance of text-based task, we utilize…

计算机视觉与模式识别 · 计算机科学 2021-10-25 Yusuf Tas , Piotr Koniusz

Embodied conversational agents (ECAs) are increasingly more realistic and capable of dynamic conversations. In online surveys, anthropomorphic agents could help address issues like careless responding and satisficing, which originate from…

人机交互 · 计算机科学 2025-08-05 Matus Krajcovic , Peter Demcak , Eduard Kuric

Discourse parsing is an important task useful for NLU applications such as summarization, machine comprehension, and emotion recognition. The current discourse parsing datasets based on conversations consists of written English dialogues…

计算与语言 · 计算机科学 2025-06-11 Divyaksh Shukla , Ritesh Baviskar , Dwijesh Gohil , Aniket Tiwari , Atul Shree , Ashutosh Modi

We present SDialog, an MIT-licensed open-source Python toolkit that unifies dialog generation, evaluation and mechanistic interpretability into a single end-to-end framework for building and analyzing LLM-based conversational agents. Built…

We present Empathic Prompting, a novel framework for multimodal human-AI interaction that enriches Large Language Model (LLM) conversations with implicit non-verbal context. The system integrates a commercial facial expression recognition…