中文
相关论文

相关论文: Beyond Point-Wise Matching: Structural Representat…

200 篇论文

Text-to-image generation powered by Diffusion Transformers (DiTs) has made remarkable strides, yet remote sensing (RS) synthesis lags behind due to two barriers: the absence of a domain-specialized DiT prior and the prohibitive cost of…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Bingxuan Zhao , Qing Zhou , Chuang Yang , Qi Wang

Referring image segmentation aims to produce a pixel-level mask for the image region described by a natural-language expression. Although pretrained vision-language models have improved semantic grounding, many existing methods still rely…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Alaa Dalaq , Muzammil Behzad

The goal of this paper is to optimize the training process of diffusion-based text-to-speech models. While recent studies have achieved remarkable advancements, their training demands substantial time and computational costs, largely due to…

音频与语音处理 · 电气工程与系统科学 2025-06-02 Jeongsoo Choi , Zhikang Niu , Ji-Hoon Kim , Chunhui Wang , Joon Son Chung , Xie Chen

We identify a core failure mode that occurs when using the usual linear interpolation on rotary positional embeddings (RoPE) for mixed-resolution denoising with Diffusion Transformers. When tokens from different spatial grids are mixed, the…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Haoyu Wu , Jingyi Xu , Qiaomu Miao , Dimitris Samaras , Hieu Le

Efficient training strategies for large-scale diffusion models have recently emphasized the importance of improving discriminative feature representations in these models. A central line of work in this direction is representation alignment…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Junno Yun , Yaşar Utku Alçalar , Mehmet Akçakaya

Protein inverse folding is a fundamental problem in bioinformatics, aiming to recover the amino acid sequences from a given protein backbone structure. Despite the success of existing methods, they struggle to fully capture the intricate…

机器学习 · 计算机科学 2024-12-13 Chenglin Wang , Yucheng Zhou , Zijie Zhai , Jianbing Shen , Kai Zhang

Recent works such as REPA have shown that guiding diffusion models with external semantic features (e.g., DINO) can significantly accelerate the training of diffusion transformers (DiTs). However, the use of pretrained external features as…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Lingchen Sun , Rongyuan Wu , Zhengqiang Zhang , Ruibin Li , Yujing Sun , Shuaizheng Liu , Lei Zhang

Enabling VLA models to predict environmental dynamics, known as world modeling, has been recognized as essential for improving robotic reasoning and generalization. However, current approaches face two main issues: 1. The training objective…

机器人学 · 计算机科学 2026-02-20 Han Zhao , Jingbo Wang , Wenxuan Song , Shuai Chen , Yang Liu , Yan Wang , Haoang Li , Donglin Wang

Scientific measurements are often bottlenecked by suboptimal conditions, whether that be noise, incomplete spatial coverage, or limited resolution, rendering accurate field reconstruction a difficult task. We introduce LatentPDE, a latent…

机器学习 · 计算机科学 2026-04-28 Valerie Tsao , Nathaniel Chaney , Manolis Veveakis

Diffusion models have revolutionized text-to-image (T2I) synthesis, producing high-quality, photorealistic images. However, they still struggle to properly render the spatial relationships described in text prompts. To address the lack of…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Andrea Rigo , Luca Stornaiuolo , Mauro Martino , Bruno Lepri , Nicu Sebe

Recent video diffusion models (VDMs) synthesize visually convincing clips, yet still drop entities, mis-bind attributes, and weaken the interactions specified in the prompt. Representation-alignment objectives such as VideoREPA and MoAlign…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Jiesong Lian , Zixiang Zhou , Ruizhe Zhong , Yuan Zhou , Qinglin Lu , Rui Wang , Long Hu , Yixue Hao , Baoru Huang

Modern ultra-high-resolution image synthesis relies heavily on the robust generative capacity of large-scale pre-trained Latent Diffusion Models (LDMs). While recent representation alignment methods have proven effective by distilling…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Jinjin Zhang , Xiefan Guo , Di Huang

Cross-modal alignment is an effective approach to improving visual classification. Existing studies typically enforce a one-step mapping that uses deep neural networks to project the visual features to mimic the distribution of textual…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Zixuan Li , Lei Meng , Guoqing Chao , Wei Wu , Xiaoshuo Yan , Yimeng Yang , Zhuang Qi , Xiangxu Meng

Predicting satellite imagery requires a balance between structural accuracy and textural detail. Standard deterministic methods like PredRNN or SimVP minimize pixel-based errors but suffer from the "regression to the mean" problem,…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Kursat Komurcu , Linas Petkevicius

Recent advances in self-supervised visual representation learning have demonstrated the effectiveness of predictive latent-space objectives for learning transferable features. In particular, Image-based Joint-Embedding Predictive…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Xiangteng He , Shunsuke Sakai , Shivam Chandhok , Sara Beery , Kun Yuan , Nicolas Padoy , Tatsuhito Hasegawa , Leonid Sigal

Remote sensing change detection is often challenged by spatial misalignment between bi-temporal images, especially when acquisitions are separated by long seasonal or multi-year gaps. While modern convolutional and transformer-based models…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Seyedehanita Madani , Vishal M. Patel

In recent years, the development of diffusion models has led to significant progress in image and video generation tasks, with pre-trained models like the Stable Diffusion series playing a crucial role. Inspired by model pruning which…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Teng Hu , Jiangning Zhang , Ran Yi , Hongrui Huang , Yabiao Wang , Lizhuang Ma

Latent diffusion models (LDMs) dominate high-quality image generation, yet integrating representation learning with generative modeling remains a challenge. We introduce a novel generative image modeling framework that seamlessly bridges…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Theodoros Kouzelis , Efstathios Karypidis , Ioannis Kakogeorgiou , Spyros Gidaris , Nikos Komodakis

Joint Embedding Predictive Architectures (JEPA) have emerged as a powerful framework for learning general-purpose representations. However, these models often lack interpretability and suffer from inefficiencies due to dense embedding…

机器学习 · 计算机科学 2025-04-24 Max Hartman , Lav Varshney

Diffusion language models (DLMs) have recently demonstrated capabilities that complement standard autoregressive (AR) models, particularly in non-sequential generation and bidirectional editing. Although recent work has shown that…

机器学习 · 计算机科学 2026-05-11 Fred Zhangzhi Peng , Alexis Fox , Anru R. Zhang , Alexander Tong