English
Related papers

Related papers: Maria: A Visual Experience Powered Conversational …

200 papers

Visual dialog is a challenging vision-language task, where a dialog agent needs to answer a series of questions through reasoning on the image content and dialog history. Prior work has mostly focused on various attention mechanisms to…

Computer Vision and Pattern Recognition · Computer Science 2020-11-03 Yue Wang , Shafiq Joty , Michael R. Lyu , Irwin King , Caiming Xiong , Steven C. H. Hoi

We build a virtual agent for learning language in a 2D maze-like world. The agent sees images of the surrounding environment, listens to a virtual teacher, and takes actions to receive rewards. It interactively learns the teacher's language…

Computation and Language · Computer Science 2018-08-15 Haonan Yu , Haichao Zhang , Wei Xu

Conversational agents ("bots") are beginning to be widely used in conversational interfaces. To design a system that is capable of emulating human-like interactions, a conversational layer that can serve as a fabric for chat-like…

Artificial Intelligence · Computer Science 2016-06-23 Abhay Prakash , Chris Brockett , Puneet Agrawal

Language interfaces with many other cognitive domains. This paper explores how interactions at these interfaces can be studied with deep learning methods, focusing on the relation between language emergence and visual perception. To model…

Social and Information Networks · Computer Science 2023-01-11 Xenia Ohmer , Michael Marino , Michael Franke , Peter König

Visual metaphors are powerful rhetorical devices used to persuade or communicate creative ideas through images. Similar to linguistic metaphors, they convey meaning implicitly through symbolism and juxtaposition of the symbols. We propose a…

Computation and Language · Computer Science 2023-07-17 Tuhin Chakrabarty , Arkadiy Saakyan , Olivia Winn , Artemis Panagopoulou , Yue Yang , Marianna Apidianaki , Smaranda Muresan

Current technological advances open up new opportunities for bringing human-machine interaction to a new level of human-centered cooperation. In this context, a key issue is the semantic understanding of the environment in order to enable…

Robotics · Computer Science 2022-11-08 Thorsten Hempel , Marc-André Fiedler , Aly Khalifa , Ayoub Al-Hamadi , Laslo Dinges

In this paper, we present a methodology for the development of embodied conversational agents for social virtual worlds. The agents provide multimodal communication with their users in which speech interaction is included. Our proposal…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-29 D. Griol , A. Sanchis , J. M. Molina , Z. Callejas

RITA presents a high-quality real-time interactive framework built upon generative models, designed with practical applications in mind. Our framework enables the transformation of user-uploaded photos into digital avatars that can engage…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Wuxinlin Cheng , Cheng Wan , Yupeng Cao , Sihan Chen

Vision-language models (VLMs) have shown to be effective at image retrieval based on simple text queries, but text-image retrieval based on conversational input remains a challenge. Consequently, if we want to use VLMs for reference…

Computation and Language · Computer Science 2023-09-26 Bram Willemsen , Livia Qian , Gabriel Skantze

In this work, we propose a deep neural architecture that uses an attention mechanism which utilizes region based image features, the natural language question asked, and semantic knowledge extracted from the regions of an image to produce…

Computation and Language · Computer Science 2021-04-06 Tasmia Tasrin , Md Sultan Al Nahian , Brent Harrison

We are increasingly surrounded by artificially intelligent technology that takes decisions and executes actions on our behalf. This creates a pressing need for general means to communicate with, instruct and guide artificial agents, with…

Today's image generation systems are capable of producing realistic and high-quality images. However, user prompts often contain ambiguities, making it difficult for these systems to interpret users' potential intentions. Consequently,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Yuheng Feng , Yangfan He , Yinghui Xia , Tianyu Shi , Jun Wang , Jinsong Yang

The rapid advances in Foundation Models and agentic Artificial Intelligence are transforming multimedia analytics by enabling richer, more sophisticated interactions between humans and analytical systems. Existing conceptual models for…

Multimedia · Computer Science 2025-04-11 Marcel Worring , Jan Zahálka , Stef van den Elzen , Maximilian T. Fischer , Daniel A. Keim

Human language is grounded on multimodal knowledge including visual knowledge like colors, sizes, and shapes. However, current large-scale pre-trained language models rely on text-only self-supervised training with massive text data, which…

Computation and Language · Computer Science 2023-02-28 Weizhi Wang , Li Dong , Hao Cheng , Haoyu Song , Xiaodong Liu , Xifeng Yan , Jianfeng Gao , Furu Wei

In this short paper, we present work evaluating an AI agent's understanding of spoken conversations about data visualizations in an online meeting scenario. There is growing interest in the development of AI-assistants that support…

Human-Computer Interaction · Computer Science 2025-10-07 Rizul Sharma , Tianyu Jiang , Seokki Lee , Jillian Aurisano

The dialogue experience with conversational agents can be greatly enhanced with multimodal and immersive interactions in virtual reality. In this work, we present an open-source architecture with the goal of simplifying the development of…

Artificial Intelligence · Computer Science 2023-08-08 Michele Yin , Gabriel Roccabruna , Abhinav Azad , Giuseppe Riccardi

Visual Dialog is a vision-language task that requires an AI agent to engage in a conversation with humans grounded in an image. It remains a challenging task since it requires the agent to fully understand a given question before making an…

Computation and Language · Computer Science 2019-12-19 Feilong Chen , Fandong Meng , Jiaming Xu , Peng Li , Bo Xu , Jie Zhou

Neural network-based systems can now learn to locate the referents of words and phrases in images, answer questions about visual scenes, and execute symbolic instructions as first-person actors in partially-observable worlds. To achieve…

Computation and Language · Computer Science 2019-10-02 Felix Hill , Stephen Clark , Karl Moritz Hermann , Phil Blunsom

Creating meaningful visual narratives through human-AI collaboration requires understanding how text-image intertextuality emerges when textual intentions meet AI-generated visuals. We conducted a three-phase qualitative study with 15…

Human-Computer Interaction · Computer Science 2025-11-06 Mengyao Guo , Kexin Nie , Ze Gao , Black Sun , Xueyang Wang , Jinda Han , Xingting Wu

Bioscientists frequently seek to visualize the biological systems they have empirically characterized and reported in the literature. Realizing such visualizations requires biological structure modeling, an inherently complex process that…

Human-Computer Interaction · Computer Science 2026-05-20 Donggang Jia , Yunhai Wang , Ivan Viola