English
Related papers

Related papers: Align and Attend: Multimodal Summarization with Du…

200 papers

This tutorial explores recent advancements in multimodal pretrained and large models, capable of integrating and processing diverse data forms such as text, images, audio, and video. Participants will gain an understanding of the…

Computation and Language · Computer Science 2024-10-10 Soyeon Caren Han , Feiqi Cao , Josiah Poon , Roberto Navigli

Abstractive multi document summarization has evolved as a task through the basic sequence to sequence approaches to transformer and graph based techniques. Each of these approaches has primarily focused on the issues of multi document…

Computation and Language · Computer Science 2022-05-10 Aiswarya Sankar , Ankit Chadha

Multimodal models integrating natural language and visual information have substantially improved generalization of representation models. However, their effectiveness significantly declines in real-world situations where certain modalities…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Jiajun Chen , Sai Cheng , Yutao Yuan , Yirui Zhang , Haitao Yuan , Peng Peng , Yi Zhong

The use of multi-modal data such as the combination of whole slide images (WSIs) and gene expression data for survival analysis can lead to more accurate survival predictions. Previous multi-modal survival models are not able to efficiently…

Computer Vision and Pattern Recognition · Computer Science 2021-12-09 Ruoqi Wang , Ziwang Huang , Haitao Wang , Hejun Wu

Audio-visual speech recognition (AVSR) system is thought to be one of the most promising solutions for robust speech recognition, especially in noisy environment. In this paper, we propose a novel multimodal attention based method for…

Computation and Language · Computer Science 2019-04-24 Pan Zhou , Wenwen Yang , Wei Chen , Yanfeng Wang , Jia Jia

Multi-modal remote sensing imagery provides complementary observations of the same geographic scene, yet such observations are frequently incomplete in practice. Existing cross-modal translation methods treat each modality pair as an…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Haoyang Chen , Jing Zhang , Hebaixu Wang , Shiqin Wang , Pohsun Huang , Jiayuan Li , Haonan Guo , Di Wang , Zheng Wang , Bo Du

Managing fluid balance in dialysis patients is crucial, as improper management can lead to severe complications. In this paper, we propose a multimodal approach that integrates visual features from lung ultrasound images with clinical data…

Image and Video Processing · Electrical Eng. & Systems 2024-10-04 Tianqi Yang , Nantheera Anantrasirichai , Oktay Karakuş , Marco Allinovi , Alin Achim

Video summarization methods are usually classified into shot-level or frame-level methods, which are individually used in a general way. This paper investigates the underlying complementarity between the frame-level and shot-level methods,…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Yubo An , Shenghui Zhao , Guoqiang Zhang

Extreme Multimodal Summarization with Multimodal Output (XMSMO) becomes an attractive summarization approach by integrating various types of information to create extremely concise yet informative summaries for individual modalities.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Sicheng Liu , Lintao Wang , Xiaogang Zhu , Xuequan Lu , Zhiyong Wang , Kun Hu

Abstractive Speech Summarization (SSum) aims to generate human-like text summaries from spoken content. It encounters difficulties in handling long speech input and capturing the intricate cross-modal mapping between long speech inputs and…

Computation and Language · Computer Science 2024-07-03 Hengchao Shang , Zongyao Li , Jiaxin Guo , Shaojun Li , Zhiqiang Rao , Yuanchang Luo , Daimeng Wei , Hao Yang

Users often struggle with decision-making between two options (A vs B), as it usually requires time-consuming research across multiple web pages. We propose STRUM-LLM that addresses this challenge by generating attributed, structured, and…

Computation and Language · Computer Science 2024-04-01 Beliz Gunel , James B. Wendt , Jing Xie , Yichao Zhou , Nguyen Vo , Zachary Fisher , Sandeep Tata

Text summarization is a user-preference based task, i.e., for one document, users often have different priorities for summary. As a key aspect of customization in summarization, granularity is used to measure the semantic coverage between…

Computation and Language · Computer Science 2022-12-15 Ming Zhong , Yang Liu , Suyu Ge , Yuning Mao , Yizhu Jiao , Xingxing Zhang , Yichong Xu , Chenguang Zhu , Michael Zeng , Jiawei Han

Effective fusion of data from multiple modalities, such as video, speech, and text, is challenging due to the heterogeneous nature of multimodal data. In this paper, we propose adaptive fusion techniques that aim to model context from…

Computation and Language · Computer Science 2021-01-27 Gaurav Sahu , Olga Vechtomova

Modern Review Helpfulness Prediction systems are dependent upon multiple modalities, typically texts and images. Unfortunately, those contemporary approaches pay scarce attention to polish representations of cross-modal relations and tend…

Computation and Language · Computer Science 2026-05-13 Thong Nguyen , Xiaobao Wu , Anh-Tuan Luu , Cong-Duy Nguyen , Zhen Hai , Lidong Bing

Audio-to-score alignment (A2SA) is a multimodal task consisting in the alignment of audio signals to music scores. Recent literature confirms the benefits of Automatic Music Transcription (AMT) for A2SA at the frame-level. In this work, we…

Sound · Computer Science 2022-01-03 Federico Simonetta , Stavros Ntalampiras , Federico Avanzini

Multimodal representation learning is fundamentally about transforming incomparable modalities into comparable representations. While prior research primarily focused on explicitly aligning these representations through targeted learning…

Machine Learning · Computer Science 2025-06-16 Megan Tjandrasuwita , Chanakya Ekbote , Liu Ziyin , Paul Pu Liang

Summarizing texts is not a straightforward task. Before even considering text summarization, one should determine what kind of summary is expected. How much should the information be compressed? Is it relevant to reformulate or should the…

Computation and Language · Computer Science 2020-07-16 Paul Tardy , David Janiszek , Yannick Estève , Vincent Nguyen

We investigate a critical yet under-explored question in Large Vision-Language Models (LVLMs): Do LVLMs genuinely comprehend interleaved image-text in the document? Existing document understanding benchmarks often assess LVLMs using…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Haolong Yan , Kaijun Tan , Yeqing Shen , Xin Huang , Zheng Ge , Xiangyu Zhang , Si Li , Daxin Jiang

The increasing volume of video content in educational, professional, and social domains necessitates effective summarization techniques that go beyond traditional unimodal approaches. This paper proposes a behaviour-aware multimodal video…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Md Moinul Islam , Sofoklis Kakouros , Janne Heikkilä , Mourad Oussalah

We study generating abstractive summaries that are faithful and factually consistent with the given articles. A novel contrastive learning formulation is presented, which leverages both reference summaries, as positive training data, and…

Computation and Language · Computer Science 2021-09-21 Shuyang Cao , Lu Wang