English
Related papers

Related papers: Modeling Text-visual Mutual Dependency for Multi-m…

200 papers

The generation of personalized dialogue is vital to natural and human-like conversation. Typically, personalized dialogue generation models involve conditioning the generated response on the dialogue history and a representation of the…

Computation and Language · Computer Science 2021-11-23 Jing Yang Lee , Kong Aik Lee , Woon Seng Gan

This paper addresses the generation of explanations with visual examples. Given an input sample, we build a system that not only classifies it to a specific category, but also outputs linguistic explanations and a set of visual examples…

Computer Vision and Pattern Recognition · Computer Science 2019-05-21 Atsushi Kanehira , Tatsuya Harada

This work seeks the possibility of generating the human face from voice solely based on the audio-visual data without any human-labeled annotations. To this end, we propose a multi-modal learning framework that links the inference stage and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-14 Hyeong-Seok Choi , Changdae Park , Kyogu Lee

Understanding and conversing about dynamic scenes is one of the key capabilities of AI agents that navigate the environment and convey useful information to humans. Video question answering is a specific scenario of such AI-human…

Computation and Language · Computer Science 2019-08-01 Guan-Lin Chao , Abhinav Rastogi , Semih Yavuz , Dilek Hakkani-Tür , Jindong Chen , Ian Lane

To explore a more scalable path for adding multimodal capabilities to existing LLMs, this paper addresses a fundamental question: Can a unimodal LLM, relying solely on text, reason about its own informational needs and provide effective…

Computation and Language · Computer Science 2026-01-13 Sazia Tabasum Mim , Jack Morris , Manish Dhakal , Yanming Xiu , Maria Gorlatova , Yi Ding

We propose an efficient method to ground pretrained text-only language models to the visual domain, enabling them to process arbitrarily interleaved image-and-text data, and generate text interleaved with retrieved images. Our method…

Computation and Language · Computer Science 2023-06-16 Jing Yu Koh , Ruslan Salakhutdinov , Daniel Fried

The improved competence of generative models can help building multi-modal virtual assistants that leverage modalities beyond language. By observing humans performing multi-step tasks, one can build assistants that have situational…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Pha Nguyen , Sailik Sengupta , Girik Malik , Arshit Gupta , Bonan Min

The popularity of image sharing on social media and the engagement it creates between users reflects the important role that visual context plays in everyday conversations. We present a novel task, Image-Grounded Conversations (IGC), in…

Computation and Language · Computer Science 2017-04-21 Nasrin Mostafazadeh , Chris Brockett , Bill Dolan , Michel Galley , Jianfeng Gao , Georgios P. Spithourakis , Lucy Vanderwende

Nowadays, the current neural network models of dialogue generation(chatbots) show great promise for generating answers for chatty agents. But they are short-sighted in that they predict utterances one at a time while disregarding their…

Computation and Language · Computer Science 2023-01-19 Jabri Ismail , Aboulbichr Ahmed , El ouaazizi Aziza

A significant ``modality gap" exists between the abundance of text-only data and the increasing power of multimodal models. This work systematically investigates whether images generated on-the-fly by Text-to-Image (T2I) models can serve as…

Multimedia · Computer Science 2026-03-04 Yuesheng Huang , Peng Zhang , Xiaoxin Wu , Riliang Liu , Jiaqi Liang

Multi-modal reasoning plays a vital role in bridging the gap between textual and visual information, enabling a deeper understanding of the context. This paper presents the Feature Swapping Multi-modal Reasoning (FSMR) model, designed to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Shuang Li , Jiahua Wang , Lijie Wen

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody expressiveness but overlooks two key issues: 1) Multiscale…

Multimedia · Computer Science 2025-01-03 Yuan Zhao , Rui Liu , Gaoxiang Cong

Web-based educational videos offer flexible learning opportunities and are becoming increasingly popular. However, improving user engagement and knowledge retention remains a challenge. Automatically generated questions can activate…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Markos Stamatakis , Joshua Berger , Christian Wartena , Ralph Ewerth , Anett Hoppe

AI models are increasingly required to be multimodal, integrating disparate input streams into a coherent state representation on which subsequent behaviors and actions can be based. This paper seeks to understand how such models behave…

Computation and Language · Computer Science 2025-07-03 Tianze Hua , Tian Yun , Ellie Pavlick

Multimodal search-based dialogue is a challenging new task: It extends visually grounded question answering systems into multi-turn conversations with access to an external database. We address this new challenge by learning a neural…

Computation and Language · Computer Science 2018-11-22 Shubham Agarwal , Ondrej Dusek , Ioannis Konstas , Verena Rieser

We present a vision and language model named MultiModal-GPT to conduct multi-round dialogue with humans. MultiModal-GPT can follow various instructions from humans, such as generating a detailed caption, counting the number of interested…

Computer Vision and Pattern Recognition · Computer Science 2023-06-14 Tao Gong , Chengqi Lyu , Shilong Zhang , Yudong Wang , Miao Zheng , Qian Zhao , Kuikun Liu , Wenwei Zhang , Ping Luo , Kai Chen

The ability to envisage the visual of a talking face based just on hearing a voice is a unique human capability. There have been a number of works that have solved for this ability recently. We differ from these approaches by enabling a…

Computer Vision and Pattern Recognition · Computer Science 2020-11-24 Ravindra Yadav , Ashish Sardana , Vinay P Namboodiri , Rajesh M Hegde

Cognitively plausible visual dialogue models should keep a mental scoreboard of shared established facts in the dialogue context. We propose a theory-based evaluation method for investigating to what degree models pretrained on the VisDial…

Computation and Language · Computer Science 2025-02-26 Brielen Madureira , David Schlangen

Traditionally, text generation models take in a sequence of text as input, and iteratively generate the next most probable word using pre-trained parameters. In this work, we propose the architecture to use images instead of text as the…

Computation and Language · Computer Science 2021-06-08 Jing Jiang
‹ Prev 1 8 9 10 Next ›