English
Related papers

Related papers: ExtraVAR: Stage-Aware RoPE Remapping for Resolutio…

200 papers

Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models are trained so that sources predicted from a mixture at an…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-08 Kohei Saijo , Gordon Wichern , François G. Germain , Zexu Pan , Jonathan Le Roux

Transformer based diffusion and vision-language models have achieved remarkable success; yet, efficiently removing undesirable or sensitive information without retraining remains a central challenge for model safety and compliance. We…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Ravi Ranjan , Utkarsh Grover , Xiaomin Lin , Agoritsa Polyzou

Camera-conditioned video generation requires positional encoding that remains reliable under changes in camera motion, lens configuration, and scene structure. However, existing attention-level camera encodings either provide ray-only…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Seonghyun Jin , Youngmin Kim , Sunwoo Park , Jong Chul Ye

Despite recent advances in Open-Vocabulary Semantic Segmentation (OVSS), existing training-free methods face several limitations: use of computationally expensive affinity refinement strategies, ineffective fusion of transformer attention…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Kunal Mahatha , Jose Dolz , Christian Desrosiers

Media compression standards have reached a plateau in terms of the rate-distortion-complexity trade-off, limiting the ability to offload expensive AI perception to the cloud in applications like robotics, wearables, and remote sensing.…

Image and Video Processing · Electrical Eng. & Systems 2026-05-29 Dan Jacobellis , Neeraja J. Yadwadkar

The vector autoregressive (VAR) model has been widely used for modeling temporal dependence in a multivariate time series. For large (and even moderate) dimensions, the number of AR coefficients can be prohibitively large, resulting in…

Applications · Statistics 2013-10-21 Richard A. Davis , Pengfei Zang , Tian Zheng

Self-supervised learning methods for computer vision have demonstrated the effectiveness of pre-training feature representations, resulting in well-generalizing Deep Neural Networks, even if the annotated data are limited. However,…

Computer Vision and Pattern Recognition · Computer Science 2021-08-25 Dmitrii Shubin , Danny Eytan , Sebastian D. Goodfellow

Knowledge distillation is a long-established technique for knowledge transfer, and has regained attention in the context of the recent emergence of large vision-language models (VLMs). However, vision-language knowledge distillation often…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Guiming Cao , Yuming Ou

Pre-trained Transformers often exhibit over-confidence in source patterns and difficulty in forming new target-domain patterns during fine-tuning. We formalize the mechanism of output saturation leading to gradient suppression through…

Machine Learning · Computer Science 2025-11-04 Wang Zixian

Generative classifiers, which leverage conditional generative models for classification, have recently demonstrated desirable properties such as robustness to distribution shifts. However, recent progress in this area has been largely…

Machine Learning · Computer Science 2026-03-24 Yi-Chung Chen , David I. Inouye , Jing Gao

High-frequency displays are gaining immense popularity because of their increasing use in video games and virtual reality applications. However, the issue is that the underlying GPUs cannot continuously generate frames at this high rate --…

Graphics · Computer Science 2023-07-25 Akanksha Dixit , Yashashwee Chakrabarty , Smruti R. Sarangi

Training neural samplers directly from unnormalized densities without access to target distribution samples presents a significant challenge. A critical desideratum in these settings is achieving comprehensive mode coverage, ensuring the…

Machine Learning · Computer Science 2025-05-27 Chenguang Wang , Xiaoyu Zhang , Kaiyuan Cui , Weichen Zhao , Yongtao Guan , Tianshu Yu

Reinforcement learning has recently been explored to improve text-to-image generation, yet applying existing GRPO algorithms to autoregressive (AR) image models remains challenging. The instability of the training process easily disrupts…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Xiaoxiao Ma , Haibo Qiu , Guohui Zhang , Zhixiong Zeng , Siqi Yang , Lin Ma , Feng Zhao

Retrieval-augmented generation (RAG) has been extensively employed to mitigate hallucinations in large language models (LLMs). However, existing methods for multi-hop reasoning tasks often lack global planning, increasing the risk of…

Computation and Language · Computer Science 2025-11-14 Yijie Zhu , Haojie Zhou , Wanting Hong , Tailin Liu , Ning Wang

The evolution of Diffusion Models has dramatically improved image generation quality, making it increasingly difficult to differentiate between real and generated images. This development, while impressive, also raises significant privacy…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Yunpeng Luo , Junlong Du , Ke Yan , Shouhong Ding

We propose ExtraNeRF, a novel method for extrapolating the range of views handled by a Neural Radiance Field (NeRF). Our main idea is to leverage NeRFs to model scene-specific, fine-grained details, while capitalizing on diffusion models to…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Meng-Li Shih , Wei-Chiu Ma , Lorenzo Boyice , Aleksander Holynski , Forrester Cole , Brian L. Curless , Janne Kontkanen

Real-world image super-resolution (SR) is a challenging image translation problem. Low-resolution (LR) images are often generated by various unknown transformations rather than by applying simple bilinear down-sampling on high-resolution…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Xin Ma , Yi Li , Huaibo Huang , Mandi Luo , Ran He

Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity problem: pretrained video diffusion transformers already…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Tianshuo Xu , Yichen Xie , Depu Meng , Chensheng Peng , Quentin Herau , Bo Jiang , Yihan Hu , Wei Zhan

Fine-tuning pre-trained generative models with Reinforcement Learning (RL) has emerged as an effective approach for aligning outputs more closely with nuanced human preferences. In this paper, we investigate the application of Group…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Matteo Gallici , Haitz Sáez de Ocáriz Borde

Recent progress in diffusion-based audio generation and restoration has substantially improved performance across heterogeneous conditioning regimes, including text-conditioned audio generation and audio-conditioned super-resolution.…

Sound · Computer Science 2026-05-07 Xuanhao Zhang , Chang Li