English
Related papers

Related papers: SAMA: Factorized Semantic Anchoring and Motion Ali…

200 papers

Exploring open-vocabulary video action recognition is a promising venture, which aims to recognize previously unseen actions within any arbitrary set of categories. Existing methods typically adapt pretrained image-text models to the video…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Chengyou Jia , Minnan Luo , Xiaojun Chang , Zhuohang Dang , Mingfei Han , Mengmeng Wang , Guang Dai , Sizhe Dang , Jingdong Wang

Due to the challenges of manually collecting accurate editing data, existing datasets are typically constructed using various automated methods, leading to noisy supervision signals caused by the mismatch between editing instructions and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Ming Li , Xin Gu , Fan Chen , Xiaoying Xing , Longyin Wen , Chen Chen , Sijie Zhu

Video-based pretraining offers immense potential for learning strong visual representations on an unprecedented scale. Recently, masked video modeling methods have shown promising scalability, yet fall short in capturing higher-level…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Mohammadreza Salehi , Michael Dorkenwald , Fida Mohammad Thoker , Efstratios Gavves , Cees G. M. Snoek , Yuki M. Asano

Segment Anything (SAM) has recently pushed the boundaries of segmentation by demonstrating zero-shot generalization and flexible prompting after training on over one billion masks. Despite this, its mask prediction accuracy often falls…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Zezhong Fan , Xiaohan Li , Topojoy Biswas , Kaushiki Nag , Kannan Achan

Video summarization helps turn long videos into clear, concise representations that are easier to review, document, and analyze, especially in high-stakes domains like surgical training. Prior work has progressed from using basic visual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Shreya Rajpal , Michal Golovanevsky , Carsten Eickhoff

Improper exposure often leads to severe loss of details, color distortion, and reduced contrast. Exposure correction still faces two critical challenges: (1) the ignorance of object-wise regional semantic information causes the color shift…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Puzhen Wu , Han Weng , Quan Zheng , Yi Zhan , Hewei Wang , Yiming Li , Jiahui Han , Rui Xu

Training-free video editing (VE) models tend to fall back on gender stereotypes when rendering profession-related prompts. We propose \textbf{FAME} for \textit{Fairness-aware Attention-modulated Video Editing} that mitigates…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Zhangkai Wu , Xuhui Fan , Zhongyuan Xie , Kaize Shi , Zhidong Li , Longbing Cao

Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approaches typically concatenate multiple videos into a single…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yue Zhang , Liqiang Jing , Jia Li , Yapeng Tian , Xinya Du , Yunhui Guo , Vibhav Gogate

Previous Sign Language Translation (SLT) methods achieve superior performance by relying on gloss annotations. However, labeling high-quality glosses is a labor-intensive task, which limits the further development of SLT. Although some…

Computation and Language · Computer Science 2024-03-20 Zhigang Chen , Benjia Zhou , Jun Li , Jun Wan , Zhen Lei , Ning Jiang , Quan Lu , Guoqing Zhao

Vision-language models (VLMs) like CLIP excel in zero-shot learning by aligning image and text representations through contrastive pretraining. Existing approaches to unsupervised adaptation (UA) for fine-grained classification with VLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Eman Ali , Sathira Silva , Chetan Arora , Muhammad Haris Khan

Large vision-language models (LVLMs) have achieved impressive results in visual question-answering and reasoning tasks through vision instruction tuning on specific datasets. However, there remains significant room for improvement in…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Xiyao Wang , Jiuhai Chen , Zhaoyang Wang , Yuhang Zhou , Yiyang Zhou , Huaxiu Yao , Tianyi Zhou , Tom Goldstein , Parminder Bhatia , Furong Huang , Cao Xiao

We address the task of zero-shot video classification for extremely fine-grained actions (e.g., Windmill Dunk in basketball), where no video examples or temporal annotations are available for unseen classes. While image-language models…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Amir Aghdam , Vincent Tao Hu , Björn Ommer

Research on Multi-modal Large Language Models (MLLMs) towards the multi-image cross-modal instruction has received increasing attention and made significant progress, particularly in scenarios involving closely resembling images (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Tao Wu , Mengze Li , Jingyuan Chen , Wei Ji , Wang Lin , Jinyang Gao , Kun Kuang , Zhou Zhao , Fei Wu

Text-guided image editing has been allowing users to transform and synthesize images through natural language instructions, offering considerable flexibility. However, most existing image editing models naively attempt to follow all user…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Hyunseung Kim , Chiho Choi , Srikanth Malla , Sai Prahladh Padmanabhan , Saurabh Bagchi , Joon Hee Choi

We present Video-LLaMA a multi-modal framework that empowers Large Language Models (LLMs) with the capability of understanding both visual and auditory content in the video. Video-LLaMA bootstraps cross-modal training from the frozen…

Computation and Language · Computer Science 2023-10-26 Hang Zhang , Xin Li , Lidong Bing

The landscape of publicly available vision foundation models (VFMs), such as CLIP and Segment Anything Model (SAM), is expanding rapidly. VFMs are endowed with distinct capabilities stemming from their pre-training objectives. For instance,…

Segment Anything Model (SAM) has gained significant recognition in the field of semantic segmentation due to its versatile capabilities and impressive performance. Despite its success, SAM faces two primary limitations: (1) it relies…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Yuchen Li , Li Zhang , Youwei Liang , Pengtao Xie

Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision mechanistic interpretability has been hindered by the lack…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Sonia Joseph , Praneet Suresh , Lorenz Hufe , Edward Stevinson , Robert Graham , Yash Vadi , Danilo Bzdok , Sebastian Lapuschkin , Lee Sharkey , Blake Aaron Richards

As the most essential property in a video, motion information is critical to a robust and generalized video representation. To inject motion dynamics, recent works have adopted frame difference as the source of motion information in video…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Minghao Zhu , Xiao Lin , Ronghao Dang , Chengju Liu , Qijun Chen

Recent advancements in language-model-based video understanding have been progressing at a remarkable pace, spurred by the introduction of Large Language Models (LLMs). However, the focus of prior research has been predominantly on devising…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Yizhou Wang , Ruiyi Zhang , Haoliang Wang , Uttaran Bhattacharya , Yun Fu , Gang Wu