English
Related papers

Related papers: Cross-view Semantic Alignment for Livestreaming Pr…

200 papers

Cross-modal entity linking refers to the ability to align entities and their attributes across different modalities. While cross-modal entity linking is a fundamental skill needed for real-world applications such as multimodal code…

Computation and Language · Computer Science 2025-06-02 Iñigo Alonso , Gorka Azkune , Ander Salaberria , Jeremy Barnes , Oier Lopez de Lacalle

With the explosive growth of video data in real-world applications, a comprehensive representation of videos becomes increasingly important. In this paper, we address the problem of video scene recognition, whose goal is to learn a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Xuzheng Yu , Chen Jiang , Wei Zhang , Tian Gan , Linlin Chao , Jianan Zhao , Yuan Cheng , Qingpei Guo , Wei Chu

Large multi-modal models (LMMs) hold the potential to usher in a new era of automated visual assistance for people who are blind or low vision (BLV). Yet, these models have not been systematically evaluated on data captured by BLV users. We…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Daniela Massiceti , Camilla Longden , Agnieszka Słowik , Samuel Wills , Martin Grayson , Cecily Morrison

With the rapid growth of live streaming platforms, personalized recommendation systems have become pivotal in improving user experience and driving platform revenue. The dynamic and multimodal nature of live streaming content (e.g., visual,…

Information Retrieval · Computer Science 2025-08-22 Yalong Guan , Xiang Chen , Mingyang Wang , Xiangyu Wu , Lihao Liu , Chao Qi , Shuang Yang , Tingting Gao , Guorui Zhou , Changjian Chen

The rapid proliferation of multimodal social media content has driven research in Multimodal Conversational Stance Detection (MCSD), which aims to interpret users' attitudes toward specific targets within complex discussions. However,…

Computation and Language · Computer Science 2026-03-11 Bingbing Wang , Zhixin Bai , Zhengda Jin , Zihan Wang , Xintong Song , Jingjie Lin , Sixuan Li , Jing Li , Ruifeng Xu

The extraction of text information in videos serves as a critical step towards semantic understanding of videos. It usually involved in two steps: (1) text recognition and (2) text classification. To localize texts in videos, we can resort…

Computer Vision and Pattern Recognition · Computer Science 2022-06-07 Ye Liu , Changchong Lu , Chen Lin , Di Yin , Bo Ren

Cross-Video Reasoning (CVR) presents a significant challenge in video understanding, which requires simultaneous understanding of multiple videos to aggregate and compare information across groups of videos. Most existing video…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Jingyao Li , Jingyun Wang , Molin Tan , Haochen Wang , Cilin Yan , Likun Shi , Jiayin Cai , Xiaolong Jiang , Yao Hu

A critical gap exists between the general-purpose visual understanding of state-of-the-art physical AI models and the specialized perceptual demands of structured real-world deployment environments. We present PRISM, a 270K-sample…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Amirreza Rouhi , Parikshit Sakurikar , Satya Sai Reddy , Narsimha Menga , Anirudh Govil , Sri Harsha Chittajallu , Rajat Aggarwal , Anoop Namboodiri , Sashi Reddi

Recent video large language models (Video LLMs) often depend on costly human annotations or proprietary model APIs (e.g., GPT-4o) to produce training data, which limits their training at scale. In this paper, we explore large-scale training…

Computer Vision and Pattern Recognition · Computer Science 2025-04-23 Joya Chen , Ziyun Zeng , Yiqi Lin , Wei Li , Zejun Ma , Mike Zheng Shou

Multimedia streaming accounts for the majority of traffic in today's internet. Mechanisms like adaptive bitrate streaming control the bitrate of a stream based on the estimated bandwidth, ideally resulting in smooth playback and a good…

Multiagent Systems · Computer Science 2024-10-29 Jannis Weil , Jonas Ringsdorf , Julian Barthel , Yi-Ping Phoebe Chen , Tobias Meuser

Accurate and complete product descriptions are crucial for e-commerce, yet seller-provided information often falls short. Customer reviews offer valuable details but are laborious to sift through manually. We present PRAISE: Product Review…

Computation and Language · Computer Science 2025-06-24 Adnan Qidwai , Srija Mukhopadhyay , Prerana Khatiwada , Dan Roth , Vivek Gupta

Multimodal fake news video detection is a crucial research direction for maintaining the credibility of online information. Existing studies primarily verify content authenticity by constructing multimodal feature fusion representations or…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Hui Li , Peien Ding , Jun Li , Guoqi Ma , Zhanyu Liu , Ge Xu , Junfeng Yao , Jinsong Su

The rapid growth of e-commerce requires robust multimodal representations that capture diverse signals from user-generated listings. Existing vision-language models (VLMs) typically align titles with primary images, i.e., single-view, but…

Information Retrieval · Computer Science 2025-12-23 Xiwen Chen , Yen-Chieh Lien , Susan Liu , María Castaños , Abolfazl Razi , Xiaoting Zhao , Congzhe Su

In adaptive bitrate streaming, resolution cross-over refers to the point on the convex hull where the encoding resolution should switch to achieve better quality. Accurate cross-over prediction is crucial for streaming providers to optimize…

Multimedia · Computer Science 2025-04-03 Jingwen Zhu , Yixu Chen , Hai Wei , Sriram Sethuraman , Yongjun Wu

The natural association between visual observations and their corresponding sound provides powerful self-supervisory signals for learning video representations, which makes the ever-growing amount of online videos an attractive source of…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Sangho Lee , Jiwan Chung , Youngjae Yu , Gunhee Kim , Thomas Breuel , Gal Chechik , Yale Song

Extracting structured information from visual documents (Visual Information Extraction, VIE) is a cornerstone of business automation. While recent Multimodal Large Language Models (MLLMs) have shown promising capabilities, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Yandi Wang , Libin Zhan , Ziwei Huang , Tiancheng Luo , Yuxuan Jiang , Wang Dong , Leilei Gan , Jun Chen

Mobile app marketplaces require developers to disclose standardized content rating descriptors (CRDs) to inform users about potentially sensitive or restricted content. Ensuring the accuracy and consistency of these disclosures remains…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Dishanika Denipitiyage , Aruna Seneviratne , Suranga Seneviratne

Remote sensing scene classification (RSSC) is a critical task with diverse applications in land use and resource management. While unimodal image-based approaches show promise, they often struggle with limitations such as high intra-class…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Jinjin Cai , Kexin Meng , Baijian Yang , Gang Shao

Since the meaning representations are detailed and accurate annotations which express fine-grained sequence-level semtantics, it is usually hard to train discriminative semantic parsers via Maximum Likelihood Estimation (MLE) in an…

Computation and Language · Computer Science 2023-01-20 Shan Wu , Chunlei Xin , Bo Chen , Xianpei Han , Le Sun

Understanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act…