English
Related papers

Related papers: DCA: Diversified Co-Attention towards Informative …

200 papers

Video content is rich in semantics and has the ability to evoke various emotions in viewers. In recent years, with the rapid development of affective computing and the explosive growth of visual data, affective video content analysis (AVCA)…

Computer Vision and Pattern Recognition · Computer Science 2024-01-19 Junxiao Xue , Jie Wang , Xuecheng Wu , Qian Zhang

Cross-platform recommendation aims to improve recommendation accuracy through associating information from different platforms. Existing cross-platform recommendation approaches assume all cross-platform information to be consistent with…

Information Retrieval · Computer Science 2019-11-13 Shengze Yu , Xin Wang , Wenwu Zhu , Peng Cui , Jingdong Wang

Live commenting on video, a popular feature of live streaming platforms, enables viewers to engage with the content and share their comments, reactions, opinions, or questions with the streamer or other viewers while watching the video or…

Computer Vision and Pattern Recognition · Computer Science 2023-11-23 Julien Lalanne , Raphael Bournet , Yi Yu

Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving forward and recent…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Zhengfei Kuang , Shengqu Cai , Hao He , Yinghao Xu , Hongsheng Li , Leonidas Guibas , Gordon Wetzstein

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive multimodal…

Computer Vision and Pattern Recognition · Computer Science 2022-04-06 Guangyao Li , Yake Wei , Yapeng Tian , Chenliang Xu , Ji-Rong Wen , Di Hu

Visual question answering (VQA) has witnessed great progress since May, 2015 as a classic problem unifying visual and textual data into a system. Many enlightening VQA works explore deep into the image and question encodings and fusing…

Computer Vision and Pattern Recognition · Computer Science 2017-02-23 Yuetan Lin , Zhangyang Pang , Donghui Wang , Yueting Zhuang

Social networks can be a valuable source of information during crisis events. In particular, users can post a stream of multimodal data that can be critical for real-time humanitarian response. However, effectively extracting meaningful…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Nusrat Munia , Junfeng Zhu , Olfa Nasraoui , Abdullah-Al-Zubaer Imran

We propose Context-aware Video-text Alignment (CVA), a novel framework to address a significant challenge in video temporal grounding: achieving temporally sensitive video-text alignment that remains robust to irrelevant background context.…

Machine Learning · Computer Science 2026-03-27 Sungho Moon , Seunghun Lee , Jiwan Seo , Sunghoon Im

Multimodal controversy detection (MCD) identifies controversial content in videos and their associated user comments, to support risk management for social video platforms.Prior research frames MCD as a static representation learning task,…

Machine Learning · Computer Science 2026-05-06 Zihan Ding , Ziyuan Yang , Yi Zhang

A major challenge for video captioning is to combine audio and visual cues. Existing multi-modal fusion methods have shown encouraging results in video understanding. However, the temporal structures of multiple modalities at different…

Computation and Language · Computer Science 2018-04-17 Xin Wang , Yuan-Fang Wang , William Yang Wang

Generative AI has significantly advanced text-driven image generation, but it still faces challenges in producing outputs that consistently align with evolving user preferences and intents, particularly in multi-turn dialogue scenarios. In…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Kun Li , Jianhui Wang , Miao Zhang , Xueqian Wang

Many real-world applications involve data from multiple modalities and thus exhibit the view heterogeneity. For example, user modeling on social media might leverage both the topology of the underlying social network and the content of the…

Machine Learning · Computer Science 2021-02-16 Lecheng Zheng , Yu Cheng , Hongxia Yang , Nan Cao , Jingrui He

Live video commenting systems are an emerging feature of online video sites. Recently the Chinese video sharing platform Bilibili, has popularised a novel captioning system where user comments are displayed as streams of moving subtitles…

Computation and Language · Computer Science 2020-06-05 Hao Wu , Gareth J. F. Jones , Francois Pitie

To generate proper captions for videos, the inference needs to identify relevant concepts and pay attention to the spatial relationships between them as well as to the temporal development in the clip. Our end-to-end encoder-decoder video…

Computer Vision and Pattern Recognition · Computer Science 2022-08-22 Zohreh Ghaderi , Leonard Salewski , Hendrik P. A. Lensch

The main idea of canonical correlation analysis (CCA) is to map different views onto a common latent space with maximum correlation. We propose a deep interpretable variational canonical correlation analysis (DICCA) for multi-view learning.…

Machine Learning · Statistics 2022-03-03 Lin Qiu , Lynn Lin , Vernon M. Chinchilli

Fact-based dialogue generation is a task of generating a human-like response based on both dialogue context and factual texts. Various methods were proposed to focus on generating informative words that contain facts effectively. However,…

Computation and Language · Computer Science 2020-05-11 Ryota Tanaka , Akinobu Lee

The rapid development of diffusion models has greatly advanced AI-generated videos in terms of length and consistency recently, yet assessing AI-generated videos still remains challenging. Previous approaches have often focused on…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Jiaze Li , Haoran Xu , Shiding Zhu , Junwei He , Haozhao Wang

Fake news often involves multimedia information such as text and image to mislead readers, proliferating and expanding its influence. Most existing fake news detection methods apply the co-attention mechanism to fuse multimodal features…

Information Retrieval · Computer Science 2023-04-13 Linmei Hu , Ziwang Zhao , Weijian Qi , Xuemeng Song , Liqiang Nie

Recent advances in multi-modal vision and language tasks enable a new set of applications. In this paper, we consider the task of generating natural language fashion feedback on outfit images. We collect a unique dataset, which contains…

Machine Learning · Computer Science 2019-06-18 Gil Sadeh , Lior Fritz , Gabi Shalev , Eduard Oks

Given a question-image input, the Visual Commonsense Reasoning (VCR) model can predict an answer with the corresponding rationale, which requires inference ability from the real world. The VCR task, which calls for exploiting the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Xuejiao Tang , Wenbin Zhang