中文
相关论文

相关论文: FrequencyBooster: Full-Frequency Modeling for High…

200 篇论文

Recent data-driven image colorization methods have enabled automatic or reference-based colorization, while still suffering from unsatisfactory and inaccurate object-level color control. To address these issues, we propose a new method…

计算机视觉与模式识别 · 计算机科学 2023-08-04 Jianxin Lin , Peng Xiao , Yijun Wang , Rongju Zhang , Xiangxiang Zeng

Data-driven flow-field reconstruction typically relies on autoencoder architectures that compress high-dimensional states into low-dimensional latent representations. However, classical approaches such as variational autoencoders (VAEs)…

机器学习 · 计算机科学 2026-01-14 AmirPouya Hemmasian , Amir Barati Farimani

Existing open-source film restoration methods show limited performance compared to commercial methods due to training with low-quality synthetic data and employing noisy optical flows. In addition, high-resolution films have not been…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Rongji Xun , Junjie Yuan , Zhongjie Wang

Diffusion probabilistic models have achieved mainstream success in many generative modeling tasks, from image generation to inverse problem solving. A distinct feature of these models is that they correspond to deep hierarchical latent…

机器学习 · 计算机科学 2024-12-30 Yibo Yang , Justus C. Will , Stephan Mandt

Diffusion posterior sampling solves inverse problems by combining a pretrained diffusion prior with measurement-consistency guidance, but it often fails to recover fine details because measurement terms are applied in a manner that is…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Feng Tian , Yixuan Li , Weili Zeng , Weitian Zhang , Yichao Yan , Xiaokang Yang

We propose FrePolad: frequency-rectified point latent diffusion, a point cloud generation pipeline integrating a variational autoencoder (VAE) with a denoising diffusion probabilistic model (DDPM) for the latent distribution. FrePolad…

计算机视觉与模式识别 · 计算机科学 2024-07-15 Chenliang Zhou , Fangcheng Zhong , Param Hanji , Zhilin Guo , Kyle Fogarty , Alejandro Sztrajman , Hongyun Gao , Cengiz Oztireli

Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions frequently struggle with phone reconstruction and…

声音 · 计算机科学 2025-02-11 Tushar Dhyani , Florian Lux , Michele Mancusi , Giorgio Fabbro , Fritz Hohl , Ngoc Thang Vu

The fast algorithms in Fourier optics have invigorated multifunctional device design and advanced imaging technologies. However, the necessity for fast computations has led to limitations in the widely used conventional Fourier methods,…

应用物理 · 物理学 2024-06-25 Zhi Li , Xuhao Luo , Jing Wang , Xin Yuan , Dongdong Teng , Qiang Song , Huigao Duan

We present Fillerbuster, a unified model that completes unknown regions of a 3D scene with a multi-view latent diffusion transformer. Casual captures are often sparse and miss surrounding content behind objects or above the scene. Existing…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Ethan Weber , Norman Müller , Yash Kant , Vasu Agrawal , Michael Zollhöfer , Angjoo Kanazawa , Christian Richardt

Advancements in deep generative models such as generative adversarial networks and variational autoencoders have resulted in the ability to generate realistic images that are visually indistinguishable from real images, which raises…

图像与视频处理 · 电气工程与系统科学 2021-02-16 Tarik Dzanic , Karan Shah , Freddie Witherden

We present ImPoster, a novel algorithm for generating a target image of a 'source' subject performing a 'driving' action. The inputs to our algorithm are a single pair of a source image with the subject that we wish to edit and a driving…

计算机视觉与模式识别 · 计算机科学 2025-02-13 Divya Kothandaraman , Kuldeep Kulkarni , Sumit Shekhar , Balaji Vasan Srinivasan , Dinesh Manocha

Latent diffusion models (LDMs) achieve state-of-the-art performance across various tasks, including image generation and video synthesis. However, they generally lack robustness, a limitation that remains not fully explored in current…

机器学习 · 计算机科学 2025-06-10 Boris Martirosyan , Alexey Karmanov

Out-of-domain (OOD) robustness under domain adaptation settings, where labeled source data and unlabeled target data come from different distributions, is a key challenge in real-world applications. A common approach to improving OOD…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Ruoqi Wang , Haitao Wang , Shaojie Guo , Qiong Luo

Point cloud compression methods jointly optimize bitrates and reconstruction distortion. However, balancing compression ratio and reconstruction quality is difficult because low-frequency and high-frequency components contribute differently…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Xiaoge Zhang , Zijie Wu , Mingtao Feng , Zichen Geng , Mehwish Nasim , Saeed Anwar , Ajmal Mian

This report presents the comprehensive implementation, evaluation, and optimization of Denoising Diffusion Probabilistic Models (DDPMs) and Denoising Diffusion Implicit Models (DDIMs), which are state-of-the-art generative models. During…

计算机视觉与模式识别 · 计算机科学 2024-12-20 Jaineet Shah , Michael Gromis , Rickston Pinto

We introduce region-specific image refinement as a dedicated problem setting: given an input image and a user-specified region (e.g., a scribble mask or a bounding box), the goal is to restore fine-grained details while keeping all…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Dewei Zhou , You Li , Zongxin Yang , Yi Yang

Video captioning aims to generate natural language sentences that describe the given video accurately. Existing methods obtain favorable generation by exploring richer visual representations in encode phase or improving the decoding…

计算机视觉与模式识别 · 计算机科学 2022-12-20 Xian Zhong , Zipeng Li , Shuqin Chen , Kui Jiang , Chen Chen , Mang Ye

Ultrasound image segmentation is pivotal for clinical diagnosis, yet challenged by speckle noise and imaging artifacts. Recently, DINOv3 has shown remarkable promise in medical image segmentation with its powerful representation…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Yixuan Zhang , Qing Xu , Yue Li , Xiangjian He , Qian Zhang , Mainul Haque , Rong Qu , Wenting Duan , Zhen Chen

Although image-based virtual try-on has made considerable progress, emerging approaches still encounter challenges in producing high-fidelity and robust fitting images across diverse scenarios. These methods often struggle with issues such…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Boyuan Jiang , Xiaobin Hu , Donghao Luo , Qingdong He , Chengming Xu , Jinlong Peng , Jiangning Zhang , Chengjie Wang , Yunsheng Wu , Yanwei Fu

Large-scale generative models, such as text-to-image diffusion models, have garnered widespread attention across diverse domains due to their creative and high-fidelity image generation. Nonetheless, existing large-scale diffusion models…

计算机视觉与模式识别 · 计算机科学 2024-08-28 Younghyun Kim , Geunmin Hwang , Junyu Zhang , Eunbyung Park
‹ 上一页 1 8 9 10 下一页 ›