中文
相关论文

相关论文: Stark: Social Long-Term Multi-Modal Conversation w…

200 篇论文

We introduce Long-VITA, a simple yet effective large multi-modal model for long-context visual-language understanding tasks. It is adept at concurrently processing and analyzing modalities of image, video, and text over 4K frames or 1M…

We present a new human-human dialogue dataset - PhotoChat, the first dataset that casts light on the photo sharing behavior in onlin emessaging. PhotoChat contains 12k dialogues, each of which is paired with a user photo that is shared…

信息检索 · 计算机科学 2021-08-04 Xiaoxue Zang , Lijuan Liu , Maria Wang , Yang Song , Hao Zhang , Jindong Chen

Conversational Recommender Systems (CRSs) aim to provide personalized recommendations by interacting with users through conversations. Most existing studies of CRS focus on extracting user preferences from conversational contexts. However,…

信息检索 · 计算机科学 2025-04-28 Yibiao Wei , Jie Zou , Weikang Guo , Guoqing Wang , Xing Xu , Yang Yang

Large language models exhibit enhanced zero-shot performance on various tasks when fine-tuned with instruction-following data. Multimodal instruction-following models extend these capabilities by integrating both text and images. However,…

计算机视觉与模式识别 · 计算机科学 2024-09-18 Yupan Huang , Zaiqiao Meng , Fangyu Liu , Yixuan Su , Nigel Collier , Yutong Lu

In this paper, we introduce a new problem, Online-MMSI, where the model must perform multimodal social interaction understanding (MMSI) using only historical information. Given a recorded video and a multi-party dialogue, the AI assistant…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Xinpeng Li , Shijian Deng , Bolin Lai , Weiguo Pian , James M. Rehg , Yapeng Tian

Shared memories between two individuals strengthen their bond and are crucial for facilitating their ongoing conversations. This study aims to make long-term dialogue more engaging by leveraging these shared memories. To this end, we…

计算与语言 · 计算机科学 2025-07-24 Eunwon Kim , Chanho Park , Buru Chang

Despite the continuing efforts to improve the engagingness and consistency of chit-chat dialogue systems, the majority of current work simply focus on mimicking human-like responses, leaving understudied the aspects of modeling…

计算与语言 · 计算机科学 2020-04-14 Qian Liu , Yihong Chen , Bei Chen , Jian-Guang Lou , Zixuan Chen , Bin Zhou , Dongmei Zhang

Emotion recognition in conversation (ERC) is a crucial component in affective dialogue systems, which helps the system understand users' emotions and generate empathetic responses. However, most works focus on modeling speaker and…

计算与语言 · 计算机科学 2021-07-15 Jingwen Hu , Yuchen Liu , Jinming Zhao , Qin Jin

Multimodal chatbots have become one of the major topics for dialogue systems in both research community and industry. Recently, researchers have shed light on the multimodality of responses as well as dialogue contexts. This work explores…

计算与语言 · 计算机科学 2026-05-05 Seongbo Jang , Seonghyeon Lee , Dongha Lee , Hwanjo Yu

Text-driven person image generation is an emerging and challenging task in cross-modality image generation. Controllable person image generation promotes a wide range of applications such as digital human interaction and virtual try-on.…

计算机视觉与模式识别 · 计算机科学 2022-11-14 Kaiduo Zhang , Muyi Sun , Jianxin Sun , Binghao Zhao , Kunbo Zhang , Zhenan Sun , Tieniu Tan

Recent works in end-to-end speech-to-text translation (ST) have proposed multi-tasking methods with soft parameter sharing which leverage machine translation (MT) data via secondary encoders that map text inputs to an eventual cross-modal…

计算与语言 · 计算机科学 2023-09-28 Brian Yan , Xuankai Chang , Antonios Anastasopoulos , Yuya Fujita , Shinji Watanabe

The body movements accompanying speech aid speakers in expressing their ideas. Co-speech motion generation is one of the important approaches for synthesizing realistic avatars. Due to the intricate correspondence between speech and motion,…

多媒体 · 计算机科学 2024-08-28 Sen Wang , Jiangning Zhang , Xin Tan , Zhifeng Xie , Chengjie Wang , Lizhuang Ma

Therapeutic art activities, such as expressive drawing and painting, require the synergy between creative visual production and interactive dialogue. Recent advancements in Multimodal Large Language Models (MLLMs) have expanded the capacity…

人机交互 · 计算机科学 2026-05-12 Le Lin , Zihao Zhu , Rainbow Tin Hung Ho , Jing Liao , Yuhan Luo

Short-form video platforms integrate text, visuals, and audio into complex communicative acts, yet existing research analyzes these modalities in isolation, lacking scalable frameworks to interpret their joint contributions. This study…

多媒体 · 计算机科学 2026-01-22 Mingyue Zha , Ho-Chun Herbert Chang

The objective of this work is person-clustering in videos -- grouping characters according to their identity. Previous methods focus on the narrower task of face-clustering, and for the most part ignore other cues such as the person's…

计算机视觉与模式识别 · 计算机科学 2021-05-21 Andrew Brown , Vicky Kalogeiton , Andrew Zisserman

Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors. However, two major challenges in modeling such multimodal human language time-series data exist: 1) inherent data…

While text-based emotion recognition methods have achieved notable success, real-world dialogue systems often demand a more nuanced emotional understanding than any single modality can offer. Multimodal Emotion Recognition in Conversations…

计算与语言 · 计算机科学 2025-09-10 Chengyan Wu , Yiqiang Cai , Yang Liu , Pengxu Zhu , Yun Xue , Ziwei Gong , Julia Hirschberg , Bolei Ma

People capture photos and videos to relive and share memories of personal significance. Recently, media montages (stories) have become a popular mode of sharing these memories due to their intuitive and powerful storytelling capabilities.…

计算与语言 · 计算机科学 2022-11-09 Satwik Kottur , Seungwhan Moon , Aram H. Markosyan , Hardik Shah , Babak Damavandi , Alborz Geramifard

Social media is daily creating massive multimedia content with paired image and text, presenting the pressing need to automate the vision and language understanding for various multimodal classification tasks. Compared to the commonly…

计算与语言 · 计算机科学 2023-03-28 Chunpu Xu , Jing Li

The integration of conversational agents into our daily lives has become increasingly common, yet many of these agents cannot engage in deep interactions with humans. Despite this, there is a noticeable shortage of datasets that capture…