English
Related papers

Related papers: HLFormer: Enhancing Partially Relevant Video Retri…

200 papers

With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating video understanding exhibit significant limitations in…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Qi Wu , Quanlong Zheng , Yanhao Zhang , Junlin Xie , Jinguo Luo , Kuo Wang , Peng Liu , Qingsong Xie , Ru Zhen , Zhenyu Yang , Haonan Lu

Standard dual-encoder vision-language models that map images and text to deterministic points on a shared unit hypersphere through $\ell_2$ normalization typically expose neither \emph{aleatoric} uncertainty (cross-modal ambiguity) nor…

Machine Learning · Computer Science 2026-05-14 Mayank Nautiyal , Li Ju , Andreas Hellander , Ekta Vats , Prashant Singh

Vision-language models have achieved remarkable success in multi-modal representation learning from large-scale pairs of visual scenes and linguistic descriptions. However, they still struggle to simultaneously express two distinct types of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Daiki Yoshikawa , Takashi Matsubara

3D visual perception tasks, including 3D detection and map segmentation based on multi-camera images, are essential for autonomous driving systems. In this work, we present a new framework termed BEVFormer, which learns unified BEV…

Computer Vision and Pattern Recognition · Computer Science 2022-07-14 Zhiqi Li , Wenhai Wang , Hongyang Li , Enze Xie , Chonghao Sima , Tong Lu , Qiao Yu , Jifeng Dai

Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vision-Language Models (LVLMs) have shown early promise in…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Ye Wang , Ziheng Wang , Boshen Xu , Yang Du , Kejun Lin , Zihan Xiao , Zihao Yue , Jianzhong Ju , Liang Zhang , Dingyi Yang , Xiangnan Fang , Zewen He , Zhenbo Luo , Wenxuan Wang , Junqi Lin , Jian Luan , Qin Jin

The rapid growth of video content demands efficient and precise retrieval systems. While vision-language models (VLMs) excel in representation learning, they often struggle with adaptive, time-sensitive video retrieval. This paper…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Yicheng Duan , Xi Huang , Duo Chen

This paper studies the multimedia problem of temporal sentence grounding (TSG), which aims to accurately determine the specific video segment in an untrimmed video according to a given sentence query. Traditional TSG methods mainly follow…

Multimedia · Computer Science 2026-05-26 Xiang Fang , Daizong Liu , Pan Zhou , Zichuan Xu , Ruixuan Li

Most real-world datasets consist of a natural hierarchy between classes or an inherent label structure that is either already available or can be constructed cheaply. However, most existing representation learning methods ignore this…

Machine Learning · Computer Science 2024-12-03 Aditya Sinha , Siqi Zeng , Makoto Yamada , Han Zhao

Empowered by Large Language Models (LLMs), recent advancements in Video-based LLMs (VideoLLMs) have driven progress in various video understanding tasks. These models encode video representations through pooling or query aggregation over a…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Yuetian Weng , Mingfei Han , Haoyu He , Xiaojun Chang , Bohan Zhuang

Current video retrieval systems, especially those used in competitions, primarily focus on querying individual keyframes or images rather than encoding an entire clip or video segment. However, queries often describe an action or event over…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Quoc-Bao Nguyen-Le , Thanh-Huy Le-Nguyen

Compression-based representations (CBRs) from neural audio codecs such as EnCodec capture intricate acoustic features like pitch and timbre, while representation-learning-based representations (RLRs) from pre-trained models trained for…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-05 Orchid Chetia Phukan , Girish , Mohd Mujtaba Akhtar , Swarup Ranjan Behera , Pailla Balakrishna Reddy , Arun Balaji Buduru , Rajesh Sharma

Transvaginal ultrasound is a critical imaging modality for evaluating cervical anatomy and detecting physiological changes. However, accurate segmentation of cervical structures remains challenging due to low contrast, shadow artifacts, and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Tran Quoc Khanh Le , Nguyen Lan Vi Vu , Ha-Hieu Pham , Xuan-Loc Huynh , Tien-Huy Nguyen , Minh Huu Nhat Le , Quan Nguyen , Hien D. Nguyen

Temporal modeling is crucial for video super-resolution. Most of the video super-resolution methods adopt the optical flow or deformable convolution for explicitly motion compensation. However, such temporal modeling techniques increase the…

Computer Vision and Pattern Recognition · Computer Science 2022-04-15 Takashi Isobe , Xu Jia , Xin Tao , Changlin Li , Ruihuang Li , Yongjie Shi , Jing Mu , Huchuan Lu , Yu-Wing Tai

The issue of data sparsity poses a significant challenge to recommender systems. In response to this, algorithms that leverage side information such as review texts have been proposed. Furthermore, Cross-Domain Recommendation (CDR), which…

Information Retrieval · Computer Science 2025-03-27 Yoonhyuk Choi , Jiho Choi , Taewook Ko , Chong-Kwon Kim

Representing data in hyperbolic space can effectively capture latent hierarchical relationships. With the goal of enabling accurate classification of points in hyperbolic space while respecting their hyperbolic geometry, we introduce…

Machine Learning · Computer Science 2018-06-04 Hyunghoon Cho , Benjamin DeMeo , Jian Peng , Bonnie Berger

Diffusion models have demonstrated exceptional capabilities in image restoration, yet their application to video super-resolution (VSR) faces significant challenges in balancing fidelity with temporal consistency. Our evaluation reveals a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Xiaohui Li , Yihao Liu , Shuo Cao , Ziyan Chen , Shaobin Zhuang , Xiangyu Chen , Yinan He , Yi Wang , Yu Qiao

As important data carriers, the drastically increasing number of multimedia videos often brings many duplicate and near-duplicate videos in the top results of search. Near-duplicate video retrieval (NDVR) can cluster and filter out the…

Information Retrieval · Computer Science 2021-06-01 Hao Cheng , Ping Wang , Chun Qi

Video anomaly detection (VAD) is crucial for intelligent surveillance, but a significant challenge lies in identifying complex anomalies, which are events defined by intricate relationships and temporal dependencies among multiple entities…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Mohammad Mahdi Hemmatyar , Mahdi Jafari , Mohammad Amin Yousefi , Mohammad Reza Nemati , Mobin Azadani , Hamid Reza Rastad , Amirmohammad Akbari

Video compression is a critical component of Internet video delivery. Recent work has shown that deep learning techniques can rival or outperform human-designed algorithms, but these methods are significantly less compute and…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Mehrdad Khani , Vibhaalakshmi Sivaraman , Mohammad Alizadeh

Reconstructing high dynamic range (HDR) images from low dynamic range (LDR) bursts plays an essential role in the computational photography. Impressive progress has been achieved by learning-based algorithms which require LDR-HDR image…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Wei Jiang , Jiahao Cui , Yizheng Wu , Zhan Peng , Zhiyu Pan , Zhiguo Cao