中文
相关论文

相关论文: OpenViDial: A Large-Scale, Open-Domain Dialogue Da…

200 篇论文

Compared to traditional visual question answering, video-grounded dialogues require additional reasoning over dialogue context to answer questions in a multi-turn setting. Previous approaches to video-grounded dialogues mostly use dialogue…

人工智能 · 计算机科学 2022-12-08 Hung Le , Nancy F. Chen , Steven C. H. Hoi

Thanks to the emerging of foundation models, the large language and vision models are integrated to acquire the multimodal ability of visual captioning, question answering, etc. Although existing multimodal models present impressive…

计算机视觉与模式识别 · 计算机科学 2023-12-29 Bo Zhao , Boya Wu , Muyang He , Tiejun Huang

We devise a multimodal conversation system for dialogue utterances composed of text, image or both modalities. We leverage Auxiliary UnsuperviseD vIsual and TExtual Data (AUDITED). To improve the performance of text-based task, we utilize…

计算机视觉与模式识别 · 计算机科学 2021-10-25 Yusuf Tas , Piotr Koniusz

Large Vision-Language Models (LVLMs) have achieved significant progress in tasks like visual question answering and document understanding. However, their potential to comprehend embodied environments and navigate within them remains…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Zhaowei Wang , Hongming Zhang , Tianqing Fang , Ye Tian , Yue Yang , Kaixin Ma , Xiaoman Pan , Yangqiu Song , Dong Yu

Designed for tracking user goals in dialogues, a dialogue state tracker is an essential component in a dialogue system. However, the research of dialogue state tracking has largely been limited to unimodality, in which slots and slot values…

人工智能 · 计算机科学 2022-06-17 Hung Le , Nancy F. Chen , Steven C. H. Hoi

Current captioning datasets focus on object-centric captions, describing the visible objects in the image, e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Language models to…

计算与语言 · 计算机科学 2023-09-26 Michele Cafagna , Kees van Deemter , Albert Gatt

Recent advancements in multi-turn voice interaction models have improved user-model communication. However, while closed-source models effectively retain and recall past utterances, whether open-source models share this ability remains…

声音 · 计算机科学 2025-05-26 Heeseung Kim , Che Hyun Lee , Sangkwon Park , Jiheum Yeom , Nohil Park , Sangwon Yu , Sungroh Yoon

Human conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge…

计算与语言 · 计算机科学 2020-03-25 Meng Chen , Ruixue Liu , Lei Shen , Shaozu Yuan , Jingyan Zhou , Youzheng Wu , Xiaodong He , Bowen Zhou

Recently, research on open domain dialogue systems have attracted extensive interests of academic and industrial researchers. The goal of an open domain dialogue system is to imitate humans in conversations. Previous works on single turn…

计算与语言 · 计算机科学 2024-10-29 Wei-Nan Zhang , Yiming Cui , Kaiyan Zhang , Yifa Wang , Qingfu Zhu , Lingzhi Li , Ting Liu

Vision-language models (VLMs) have shown to be effective at image retrieval based on simple text queries, but text-image retrieval based on conversational input remains a challenge. Consequently, if we want to use VLMs for reference…

计算与语言 · 计算机科学 2023-09-26 Bram Willemsen , Livia Qian , Gabriel Skantze

The recent advance in vision-language models is largely attributed to the abundance of image-text data. We aim to replicate this success for video-language models, but there simply is not enough human-curated video-text data available. We…

计算机视觉与模式识别 · 计算机科学 2024-04-17 Yue Zhao , Long Zhao , Xingyi Zhou , Jialin Wu , Chun-Te Chu , Hui Miao , Florian Schroff , Hartwig Adam , Ting Liu , Boqing Gong , Philipp Krähenbühl , Liangzhe Yuan

Emotion Recognition in Conversation is a core component of affective computing, while current resources of sign language emotion datasets primarily focus on isolated sentences and lack conversational context. Models trained exclusively on…

计算与语言 · 计算机科学 2026-05-25 Yusong Wang , Keyu Mao , Takao Obi , Minghao Shao , Kotaro Funakoshi

Recent advances in AI-driven storytelling have enhanced video generation and story visualization. However, translating dialogue-centric scripts into coherent storyboards remains a significant challenge due to limited script detail,…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Min Zhang , Zilin Wang , Liyan Chen , Kunhong Liu , Juncong Lin

The research of knowledge-driven conversational systems is largely limited due to the lack of dialog data which consist of multi-turn conversations on multiple topics and with knowledge annotations. In this paper, we propose a Chinese…

计算与语言 · 计算机科学 2020-04-09 Hao Zhou , Chujie Zheng , Kaili Huang , Minlie Huang , Xiaoyan Zhu

We present a vision and language model named MultiModal-GPT to conduct multi-round dialogue with humans. MultiModal-GPT can follow various instructions from humans, such as generating a detailed caption, counting the number of interested…

计算机视觉与模式识别 · 计算机科学 2023-06-14 Tao Gong , Chengqi Lyu , Shilong Zhang , Yudong Wang , Miao Zheng , Qian Zhao , Kuikun Liu , Wenwei Zhang , Ping Luo , Kai Chen

This paper presents the Frames dataset (Frames is available at http://datasets.maluuba.com/Frames), a corpus of 1369 human-human dialogues with an average of 15 turns per dialogue. We developed this dataset to study the role of memory in…

This work aims to create a multimodal AI system that chats with humans and shares relevant photos. While earlier works were limited to dialogues about specific objects or scenes within images, recent works have incorporated images into…

计算与语言 · 计算机科学 2023-05-08 Min Young Lee

Human language is grounded on multimodal knowledge including visual knowledge like colors, sizes, and shapes. However, current large-scale pre-trained language models rely on text-only self-supervised training with massive text data, which…

计算与语言 · 计算机科学 2023-02-28 Weizhi Wang , Li Dong , Hao Cheng , Haoyu Song , Xiaodong Liu , Xifeng Yan , Jianfeng Gao , Furu Wei

Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent AI technologies, it is crucial to develop models that can…

Dialogue systems have been widely applied in many scenarios and are now more powerful and ubiquitous than ever before. With large neural models and massive available data, current dialogue systems have access to more knowledge than any…