中文
相关论文

相关论文: SWIFT: Sliding Window Reconstruction for Few-Shot …

200 篇论文

Recent advancements in dataset distillation have demonstrated the significant benefits of employing soft labels generated by pre-trained teacher models. In this paper, we introduce a novel perspective by emphasizing the full utilization of…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Xinyi Shang , Peng Sun , Tao Lin

Video stylization plays a key role in content creation, but it remains a challenging problem. Na\"ively applying image stylization frame-by-frame hurts temporal consistency and reduces style richness. Alternatively, training a dedicated…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Jiacong Xu , Yiqun Mei , Ke Zhang , Vishal M. Patel

Despite recent advances in Novel View Synthesis (NVS), generating high-fidelity views from single or sparse observations remains a significant challenge. Existing splatting-based approaches often produce distorted geometry due to splatting…

计算机视觉与模式识别 · 计算机科学 2025-02-19 Xiang Zhang , Yang Zhang , Lukas Mehl , Markus Gross , Christopher Schroers

Video generation primarily aims to model authentic and customized motion across frames, making understanding and controlling the motion a crucial topic. Most diffusion-based studies on video motion focus on motion customization with…

计算机视觉与模式识别 · 计算机科学 2024-11-13 Zeqi Xiao , Yifan Zhou , Shuai Yang , Xingang Pan

Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a motion-centric, gradient-based data attribution framework…

计算机视觉与模式识别 · 计算机科学 2026-01-14 Xindi Wu , Despoina Paschalidou , Jun Gao , Antonio Torralba , Laura Leal-Taixé , Olga Russakovsky , Sanja Fidler , Jonathan Lorraine

This paper proposes a Short-Window Sliding Learning framework for real-time violence detection in CCTV footages. Unlike conventional long-video training approaches, the proposed method divides videos into 1-2 second clips and applies Large…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Seoik Jung , Taekyung Song , Yangro Lee , Sungjun Lee

Following major advances in text and image generation, the video domain has surged, producing highly realistic and controllable sequences. Along with this progress, these models also raise serious concerns about misinformation, making…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Omer Ben Hayun , Roy Betser , Meir Yossef Levi , Levi Kassel , Guy Gilboa

Zero-shot sketch-based image retrieval typically asks for a trained model to be applied as is to unseen categories. In this paper, we question to argue that this setup by definition is not compatible with the inherent abstract and…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Aneeshan Sain , Ayan Kumar Bhunia , Vaishnav Potlapalli , Pinaki Nath Chowdhury , Tao Xiang , Yi-Zhe Song

Transformers have achieved the state-of-the-art performance on solving the inverse problem of Snapshot Compressive Imaging (SCI) for video, whose ill-posedness is rooted in the mixed degradation of spatial masking and temporal aliasing.…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Ping Wang , Yulun Zhang , Lishun Wang , Xin Yuan

In performative prediction, the deployment of a predictive model triggers a shift in the data distribution. As these shifts are typically unknown ahead of time, the learner needs to deploy a model to get feedback about the distribution it…

机器学习 · 计算机科学 2022-07-19 Meena Jagadeesan , Tijana Zrnic , Celestine Mendler-Dünner

Despite the recent success of neural networks in image feature learning, a major problem in the video domain is the lack of sufficient labeled data for learning to model temporal information. In this paper, we propose an unsupervised…

计算机视觉与模式识别 · 计算机科学 2016-11-29 Linchao Zhu , Zhongwen Xu , Yi Yang

The quadratic complexity of self attention in Transformer based LLMs renders long context inference prohibitively expensive. While Sliding Window Attention (SWA), the simplest sparse attention pattern, offers a linear complexity…

计算与语言 · 计算机科学 2026-03-27 Yijiong Yu , Jiale Liu , Qingyun Wu , Huazheng Wang , Ji Pei

We present an amortized framework for real-time visual attribution streaming in multimodal thinking models. When these models generate code from a screenshot or solve math problems from images, their long reasoning traces should be grounded…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Seil Kang , Woojung Han , Junhyeok Kim , Jinyeong Kim , Youngeun Kim , Seong Jae Hwang

Transformer's recent integration into style transfer leverages its proficiency in establishing long-range dependencies, albeit at the expense of attenuated local modeling. This paper introduces Strips Window Attention Transformer (S2WAT), a…

计算机视觉与模式识别 · 计算机科学 2023-12-18 Chiyu Zhang , Xiaogang Xu , Lei Wang , Zaiyan Dai , Jun Yang

In many video restoration/translation tasks, image processing operations are na\"ively extended to the video domain by processing each frame independently, disregarding the temporal connection of the video frames. This disregard for the…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Muhammad Kashif Ali , Dongjin Kim , Tae Hyun Kim

Vision foundation models achieve remarkable performance but are only available in a limited set of pre-determined sizes, forcing sub-optimal deployment choices under real-world constraints. We introduce SnapViT: Single-shot network…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Walter Simoncini , Michael Dorkenwald , Tijmen Blankevoort , Cees G. M. Snoek , Yuki M. Asano

Video compression aims to maximize reconstruction quality with minimal bitrates. Beyond standard distortion metrics, perceptual quality and temporal consistency are also critical. However, at ultra-low bitrates, traditional end-to-end…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Mingde Zhou , Zheng Chen , Yulun Zhang

Face-swapping models have been drawing attention for their compelling generation quality, but their complex architectures and loss functions often require careful tuning for successful training. We propose a new face-swapping model called…

计算机视觉与模式识别 · 计算机科学 2022-05-06 Jiseob Kim , Jihoon Lee , Byoung-Tak Zhang

Few-shot video classification aims to learn new video categories with only a few labeled examples, alleviating the burden of costly annotation in real-world applications. However, it is particularly challenging to learn a class-invariant…

计算机视觉与模式识别 · 计算机科学 2021-05-12 Songyang Zhang , Jiale Zhou , Xuming He

Makeup transfer is not only to extract the makeup style of the reference image, but also to render the makeup style to the semantic corresponding position of the target image. However, most existing methods focus on the former and ignore…

计算机视觉与模式识别 · 计算机科学 2021-12-08 Zhaoyang Sun , Yaxiong Chen , Shengwu Xiong