中文
相关论文

相关论文: Continual Text-to-Video Retrieval with Frame Fusio…

200 篇论文

Diffusion models excel in noise-to-data generation tasks, providing a mapping from a Gaussian distribution to a more complex data distribution. However they struggle to model translations between complex distributions, limiting their…

机器学习 · 计算机科学 2026-03-27 Viacheslav Vasilev , Arseny Ivanov , Nikita Gushchin , Maria Kovaleva , Alexander Korotin

Diffusion and flow matching models have significantly advanced media generation, yet their design space is well-explored, somewhat limiting further improvements. Concurrently, autoregressive (AR) models, particularly those generating…

机器学习 · 计算机科学 2025-07-01 Neta Shaul , Uriel Singer , Itai Gat , Yaron Lipman

Despite the recent success of single image-based 3D human pose and shape estimation methods, recovering temporally consistent and smooth 3D human motion from a video is still challenging. Several video-based methods have been proposed;…

计算机视觉与模式识别 · 计算机科学 2021-04-28 Hongsuk Choi , Gyeongsik Moon , Ju Yong Chang , Kyoung Mu Lee

With the rapid growth of video data, Composed Video Retrieval (CVR) has emerged as a novel paradigm in video retrieval and is receiving increasing attention from researchers. Unlike unimodal video retrieval methods, the CVR task takes a…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Zixu Li , Yupeng Hu , Zhiwei Chen , Qinlei Huang , Guozhi Qiu , Zhiheng Fu , Meng Liu

Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual question answering, cross-modal retrieval, and caption generation.…

计算机视觉与模式识别 · 计算机科学 2025-07-29 G. Thomas Hudson , Dean Slack , Thomas Winterbottom , Jamie Sterling , Chenghao Xiao , Junjie Shentu , Noura Al Moubayed

We present a method for trajectory planning for autonomous driving, learning image-based context embeddings that align with motion prediction frameworks and planning-based intention input. Within our method, a ViT encoder takes raw images…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Maitrayee Keskar , Mohan Trivedi , Ross Greer

Modern video-text retrieval frameworks basically consist of three parts: video encoder, text encoder and the similarity head. With the success on both visual and textual representation learning, transformer based encoders and fusion methods…

计算机视觉与模式识别 · 计算机科学 2022-07-22 Zijian Gao , Jingyu Liu , Weiqi Sun , Sheng Chen , Dedan Chang , Lili Zhao

The real world is dynamic, yet most image fusion methods process static frames independently, ignoring temporal correlations in videos and leading to flickering and temporal inconsistency. To address this, we propose Unified Video Fusion…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Zixiang Zhao , Haowen Bai , Bingxin Ke , Yukun Cui , Lilun Deng , Yulun Zhang , Kai Zhang , Konrad Schindler

Perceptual video quality assessment models are either frame-based or video-based, i.e., they apply spatiotemporal filtering or motion estimation to capture temporal video distortions. Despite their good performance on video quality…

图像与视频处理 · 电气工程与系统科学 2018-04-16 Christos G. Bampis , Zhi Li , Alan C. Bovik

Recently, Video-Language Models (VideoLMs) have demonstrated remarkable capabilities, offering significant potential for flexible and powerful video query systems. These models typically rely on Vision Transformers (ViTs), which process…

分布式、并行与集群计算 · 计算机科学 2025-09-11 Jinwoo Hwang , Daeun Kim , Sangyeop Lee , Yoonsung Kim , Guseul Heo , Hojoon Kim , Yunseok Jeong , Tadiwos Meaza , Eunhyeok Park , Jeongseob Ahn , Jongse Park

The VALUE (Video-And-Language Understanding Evaluation) benchmark is newly introduced to evaluate and analyze multi-modal representation learning algorithms on three video-and-language tasks: Retrieval, QA, and Captioning. The main…

计算机视觉与模式识别 · 计算机科学 2021-10-14 Minchul Shin , Jonghwan Mun , Kyoung-Woon On , Woo-Young Kang , Gunsoo Han , Eun-Sol Kim

In current text-to-video retrieval (T2VR), videos to be retrieved have been properly trimmed so that a correspondence between the videos and ad-hoc textual queries naturally exists. Note in practice that videos circulated on the Internet…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Xianke Chen , Daizong Liu , Xun Yang , Xirong Li , Jianfeng Dong , Meng Wang , Xun Wang

In the realm of multi-object tracking, the challenge of accurately capturing the spatial and temporal relationships between objects in video sequences remains a significant hurdle. This is further complicated by frequent occurrences of…

计算机视觉与模式识别 · 计算机科学 2025-01-20 Futian Wang , Fengxiang Liu , Xiao Wang

Due to the great saving of computation and memory overhead, token compression has become a research hot-spot for MLLMs and achieved remarkable progress in image-language tasks. However, for the video, existing methods still fall short of…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Shaobo Ju , Baiyang Song , Tao Chen , Jiapeng Zhang , Qiong Wu , Chao Chang , HuaiXi Wang , Yiyi Zhou , Rongrong Ji

This paper presents FluxMem, a training-free framework for efficient streaming video understanding. FluxMem adaptively compresses redundant visual memory through a hierarchical, two-stage design: (1) a Temporal Adjacency Selection (TAS)…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Yiweng Xie , Bo He , Junke Wang , Xiangyu Zheng , Ziyi Ye , Zuxuan Wu

Spatial convolutions are extensively used in numerous deep video models. It fundamentally assumes spatio-temporal invariance, i.e., using shared weights for every location in different frames. This work presents Temporally-Adaptive…

计算机视觉与模式识别 · 计算机科学 2023-08-14 Ziyuan Huang , Shiwei Zhang , Liang Pan , Zhiwu Qing , Yingya Zhang , Ziwei Liu , Marcelo H. Ang

Image animation has seen significant progress, driven by the powerful generative capabilities of diffusion models. However, maintaining appearance consistency with static input images and mitigating abrupt motion transitions in generated…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Xin Ma , Yaohui Wang , Genyun Jia , Xinyuan Chen , Tien-Tsin Wong , Cunjian Chen

Neural fields, also known as coordinate-based or implicit neural representations, have shown a remarkable capability of representing, generating, and manipulating various forms of signals. For video representations, however, mapping…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Joo Chan Lee , Daniel Rho , Jong Hwan Ko , Eunbyung Park

Composed Image Retrieval (CIR) retrieves target images using a reference image paired with modification text. Despite rapid advances, all existing methods and datasets operate at the image level -- a single reference image plus modification…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Peng Yuan , Bingyin Mei , Hui Zhang

Due to recent advances in pose-estimation methods, human motion can be extracted from a common video in the form of 3D skeleton sequences. Despite wonderful application opportunities, effective and efficient content-based access to large…

计算机视觉与模式识别 · 计算机科学 2023-10-05 Nicola Messina , Jan Sedmidubsky , Fabrizio Falchi , Tomáš Rebok
‹ 上一页 1 8 9 10 下一页 ›