English
Related papers

Related papers: Em-Garde: A Propose-Match Framework for Proactive …

200 papers

Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Shenghao Fu , Qize Yang , Yuan-Ming Li , Yi-Xing Peng , Kun-Yu Lin , Xihan Wei , Jian-Fang Hu , Xiaohua Xie , Wei-Shi Zheng

In recent years, social media users have spent significant amounts of time on short-form video platforms. As a result, established platforms in other domains, such as e-commerce, have begun introducing short-form video content to engage…

Machine Learning · Computer Science 2025-09-05 Andrii Dzhoha , Katya Mirylenka , Egor Malykh , Marco-Andrea Buchmann , Francesca Catino

The shift toward IoT-enabled, sensor-driven systems has transformed how operational data is generated, favoring continuous, real-time event streams (ES) over static event logs. This evolution presents new challenges for Streaming Process…

Video understanding with multimodal large language models (MLLMs) remains challenging due to the long token sequences of videos, which contain extensive temporal dependencies and redundant frames. Existing approaches typically treat MLLMs…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Yaolun Zhang , Ruohui Wang , Jiahao Wang , Yepeng Tang , Xuanyu Zheng , Haonan Duan , Hao Lu , Hanming Deng , Lewei Lu

Owing to powerful natural language processing and generative capabilities, large language model (LLM) agents have emerged as a promising solution for enhancing recommendation systems via user simulation. However, in the realm of video…

Multimedia · Computer Science 2025-07-04 Siran Chen , Boyu Chen , Chenyun Yu , Yuxiao Luo , Ouyang Yi , Lei Cheng , Chengxiang Zhuo , Zang Li , Yali Wang

While neural networks have excelled in video action recognition tasks, their black-box nature often obscures the understanding of their decision-making processes. Recent approaches used inherently interpretable models to analyze video…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Ning Wang , Guangming Zhu , HS Li , Liang Zhang , Syed Afaq Ali Shah , Mohammed Bennamoun

In this paper, we present a vision for a new generation of multimodal streaming systems that embed MLLMs as first-class operators, enabling real-time query processing across multiple modalities. Achieving this is non-trivial: while recent…

Event-Level Video Question Answering (EVQA) requires complex reasoning across video events to obtain the visual information needed to provide optimal answers. However, despite significant progress in model performance, few studies have…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Chenyang Lyu , Tianbo Ji , Yvette Graham , Jennifer Foster

Understanding of video creativity and content often varies among individuals, with differences in focal points and cognitive levels across different ages, experiences, and genders. There is currently a lack of research in this area, and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Minghui Wu , Chenxu Zhao , Anyang Su , Donglin Di , Tianyu Fu , Da An , Min He , Ya Gao , Meng Ma , Kun Yan , Ping Wang

Streaming video understanding requires models to robustly encode, store, and retrieve information from a continuous video stream to support accurate video question answering (VQA). Existing state-of-the-art approaches rely on key-value…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Vatsal Agarwal , Saksham Suri , Matthew Gwilliam , Pulkit Kumar , Abhinav Shrivastava

Stimulated by the sophisticated reasoning capabilities of recent Large Language Models (LLMs), a variety of strategies for bridging video modality have been devised. A prominent strategy involves Video Language Models (VideoLMs), which…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Wonkyun Kim , Changin Choi , Wonseok Lee , Wonjong Rhee

In this paper, we present a framework for the dynamic selection of the wireless channels used to deliver information-rich data streams to edge servers. The approach we propose is data-driven, where a predictor, whose output informs the…

Networking and Internet Architecture · Computer Science 2020-04-06 Sabur Baidya , Peyman Tehrani , Marco Levorato

In modern human-robot collaboration (HRC) applications, multiple perception modules jointly extract visual, auditory, and contextual cues to achieve comprehensive scene understanding, enabling the robot to provide appropriate assistance to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Dingcheng Huang , Xiaotong Zhang , Kamal Youcef-Toumi

Wireframe parsing aims to recover line segments and their junctions to form a structured geometric representation useful for downstream tasks such as Simultaneous Localization and Mapping (SLAM). Existing methods predict lines and junctions…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Chao Wang , Xuanying Li , Cheng Dai , Jinglei Feng , Yuxiang Luo , Yuqi Ouyang , Hao Qin

Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insufficient. To address this, we introduce InstructionBench, an…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Haiwan Wei , Yitian Yuan , Xiaohan Lan , Wei Ke , Lin Ma

Real-world recommendation systems commonly offer diverse content scenarios for users to interact with. Considering the enormous number of users in industrial platforms, it is infeasible to utilize a single unified recommendation model to…

Information Retrieval · Computer Science 2024-12-10 Chonggang Song , Chunxu Shen , Hao Gu , Yaoming Wu , Lingling Yi , Jie Wen , Chuan Chen

Prompt learning represents a promising method for adapting pre-trained vision-language models (VLMs) to various downstream tasks by learning a set of text embeddings. One challenge inherent to these methods is the poor generalization…

Computer Vision and Pattern Recognition · Computer Science 2024-11-18 Fangming Cui , Xun Yang , Chao Wu , Liang Xiao , Xinmei Tian

Live-streaming, as an emerging media enabling real-time interaction between authors and users, has attracted significant attention. Unlike the stable playback time of traditional TV live or the fixed content of short video, live-streaming,…

Information Retrieval · Computer Science 2025-12-09 Jiangxia Cao , Ruochen Yang , Xiang Chen , Changxin Lao , Yueyang Liu , Yusheng Huang , Yuanhao Tian , Xiangyu Wu , Shuang Yang , Zhaojie Liu , Guorui Zhou

This paper develops an edge-device collaborative Generative Semantic Communications (Gen SemCom) framework leveraging pre-trained Multi-modal/Vision Language Models (M/VLMs) for ultra-low-rate semantic communication via textual prompts. The…

Information Theory · Computer Science 2025-05-05 Mengmeng Ren , Li Qiao , Long Yang , Zhen Gao , Jian Chen , Mahdi Boloursaz Mashhadi , Pei Xiao , Rahim Tafazolli , Mehdi Bennis

Multimodal large language models (MLLMs) extend the success of language models to visual understanding, and recent efforts have sought to build unified MLLMs that support both understanding and generation. However, constructing such models…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Hanyu Wang , Jiaming Han , Ziyan Yang , Qi Zhao , Shanchuan Lin , Xiangyu Yue , Abhinav Shrivastava , Zhenheng Yang , Hao Chen