English
Related papers

Related papers: Alleviating Video-Length Effect for Micro-video Re…

200 papers

The growing trend of sharing short videos on social media platforms, where users capture and share moments from their daily lives, has led to an increase in research efforts focused on micro-video recommendations. However, conventional…

Information Retrieval · Computer Science 2025-04-07 Sanghyuck Lee , Sangkeun Park , Jaesung Lee

An effective online recommendation system should jointly capture users' long-term and short-term preferences in both users' internal behaviors (from the target recommendation task) and external behaviors (from other tasks). However, it is…

Information Retrieval · Computer Science 2021-11-29 Ruobing Xie , Yalong Wang , Rui Wang , Yuanfu Lu , Yuanhang Zou , Feng Xia , Leyu Lin

Long-form video understanding poses a significant challenge for video large language models (VideoLLMs) due to prohibitively high computational and memory demands. In this paper, we propose FlexSelect, a flexible and efficient token…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Yunzhu Zhang , Yu Lu , Tianyi Wang , Fengyun Rao , Yi Yang , Linchao Zhu

Kuaishou serving hundreds of millions of searches daily, the quality of short-video search is paramount. However, it suffers from a severe Matthew effect on long-tail queries: sparse user behavior data causes models to amplify low-quality…

Information Retrieval · Computer Science 2026-03-31 Wenyi Xu , Feiran Zhu , Songyang Li , Renzhe Zhou , Chao Zhang , Chenglei Dai , Yuren Mao , Yunjun Gao , Yi Zhang

Video large multimodal models increasingly face a scalability bottleneck: long videos produce excessively long visual-token sequences, which sharply increase memory and latency during inference. While existing compression methods are…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Kuanwei Lin , Wenhao Zhang , Ge Li

Micro-video popularity prediction (MVPP) aims to forecast the future popularity of videos on online media, which is essential for applications such as content recommendation and traffic allocation. In real-world scenarios, it is critical…

Multimedia · Computer Science 2026-04-24 Dali Wang , Yunyao Zhang , Junqing Yu , Yi-Ping Phoebe Chen , Chen Xu , Zikai Song

Short-form videos have become one of the most popular user-generated content formats nowadays. Popular short-video platforms use a simple streaming approach that preloads one or more videos in the recommendation list in advance. However,…

Multimedia · Computer Science 2026-03-25 Vu Thi Hai Yen , Duc V. Nguyen , Cao Anh Minh Huy , Truong Thu Huong

Recent advances in Large Language Models (LLMs) have opened new possibilities for recommendation systems, though current approaches such as TALLRec face challenges in explainability and cold-start scenarios. We present ExplainRec, a…

Information Retrieval · Computer Science 2025-11-20 Bo Ma , LuYao Liu , ZeHua Hu , Simon Lau

In this paper, we propose to learn temporal embeddings of video frames for complex video analysis. Large quantities of unlabeled video data can be easily obtained from the Internet. These videos possess the implicit weak label that they are…

Computer Vision and Pattern Recognition · Computer Science 2015-05-05 Vignesh Ramanathan , Kevin Tang , Greg Mori , Li Fei-Fei

Volumetric video emerges as a new attractive video paradigm in recent years since it provides an immersive and interactive 3D viewing experience with six degree-of-freedom (DoF). Unlike traditional 2D or panoramic videos, volumetric videos…

Multimedia · Computer Science 2023-08-17 Kaiyuan Hu , Haowen Yang , Yili Jin , Junhua Liu , Yongting Chen , Miao Zhang , Fangxin Wang

Multimodal Large Language Models (MLLMs) have shown promising progress in understanding and analyzing video content. However, processing long videos remains a significant challenge constrained by LLM's context size. To address this…

In recommender system, some feature directly affects whether an interaction would happen, making the happened interactions not necessarily indicate user preference. For instance, short videos are objectively easier to be finished even…

Information Retrieval · Computer Science 2022-08-29 Xiangnan He , Yang Zhang , Fuli Feng , Chonggang Song , Lingling Yi , Guohui Ling , Yongdong Zhang

With the development of multimedia systems, multimodal recommendations are playing an essential role, as they can leverage rich contexts beyond interactions. Existing methods mainly regard multimodal information as an auxiliary, using them…

Information Retrieval · Computer Science 2024-08-02 Yifan Liu , Kangning Zhang , Xiangyuan Ren , Yanhua Huang , Jiarui Jin , Yingjie Qin , Ruilong Su , Ruiwen Xu , Yong Yu , Weinan Zhang

Recently, learned video compression (LVC) is undergoing a period of rapid development. However, due to absence of large and high-quality high dynamic range (HDR) video training data, LVC on HDR video is still unexplored. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Zhaoyi Tian , Feifeng Wang , Shiwei Wang , Zihao Zhou , Yao Zhu , Liquan Shen

Recent progress in multi-modal large language models (MLLMs) has significantly advanced video understanding. However, their performance on long-form videos remains limited by computational constraints and suboptimal frame selection. We…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Wenhui Tan , Ruihua Song , Jiaze Li , Jianzhong Ju , Zhenbo Luo

Learning from demonstrations (LfD) typically relies on large amounts of action-labeled expert trajectories, which fundamentally constrains the scale of available training data. A promising alternative is to learn directly from unlabeled…

Robotics · Computer Science 2025-08-13 Haoyu Zhang , Long Cheng

Streaming video understanding with large vision-language models (VLMs) requires a compact memory that can support future reasoning over an ever-growing visual history. A common solution is to compress the key-value (KV) cache, but existing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Ailar Mahdizadeh , Puria Azadi , Muchen Li , Xiangteng He , Leonid Sigal

In recommender systems, the user-item interaction data is usually sparse and not sufficient for learning comprehensive user/item representations for recommendation. To address this problem, we propose a novel dual-bridging recommendation…

Information Retrieval · Computer Science 2019-10-17 Jingwei Ma , Jiahui Wen , Mingyang Zhong , Liangchen Liu , Chaojie Li , Weitong Chen , Yin Yang , Honghui Tu , Xue Li

The advancements in large language models (LLMs) have propelled the improvement of video understanding tasks by incorporating LLMs with visual models. However, most existing LLM-based models (e.g., VideoLLaMA, VideoChat) are constrained to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yuanbin Man , Ying Huang , Chengming Zhang , Bingzhe Li , Wei Niu , Miao Yin

Long videos, ranging from minutes to hours, present significant challenges for current Multi-modal Large Language Models (MLLMs) due to their complex events, diverse scenes, and long-range dependencies. Direct encoding of such videos is…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Zizhong Li , Haopeng Zhang , Jiawei Zhang