中文
相关论文

相关论文: SmartSpatial: Enhancing the 3D Spatial Arrangement…

200 篇论文

Multimodal Small-to-Medium sized Language Models (MSLMs) have demonstrated strong capabilities in integrating visual and textual information but still face significant limitations in visual comprehension and mathematical reasoning,…

机器学习 · 计算机科学 2026-01-27 Ashutosh Bajpai , Akshat Bhandari , Akshay Nambi , Tanmoy Chakraborty

The slow inference process of image diffusion models significantly degrades interactive user experiences. To address this, we introduce Diffusion Preview, a novel paradigm employing rapid, low-step sampling to generate preliminary outputs…

Despite their generative power, diffusion models struggle to maintain style consistency across images conditioned on the same style prompt, hindering their practical deployment in creative workflows. While several training-free methods…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Jiexuan Zhang , Yiheng Du , Qian Wang , Weiqi Li , Yu Gu , Jian Zhang

We propose SparseFusion, a sparse view 3D reconstruction approach that unifies recent advances in neural rendering and probabilistic image generation. Existing approaches typically build on neural rendering with re-projected features but…

计算机视觉与模式识别 · 计算机科学 2023-02-17 Zhizhuo Zhou , Shubham Tulsiani

This paper presents an end-to-end neural architecture based on Diffusion Models for spatial puzzle solving, particularly jigsaw puzzle and room arrangement tasks. In the latter task, for instance, the proposed system "PuzzleFusion" takes a…

人工智能 · 计算机科学 2023-10-05 Sepidehsadat Hosseini , Mohammad Amin Shabani , Saghar Irandoust , Yasutaka Furukawa

Recent advances in 3D representations, such as Neural Radiance Fields and 3D Gaussian Splatting, have greatly improved realistic scene modeling and novel-view synthesis. However, achieving controllable and consistent editing in dynamic 3D…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Kai He , Chin-Hsuan Wu , Igor Gilitschenski

Pseudo-healthy image inpainting is an essential preprocessing step for analyzing pathological brain MRI scans. Most current inpainting methods favor slice-wise 2D models for their high in-plane fidelity, but their independence across slices…

图像与视频处理 · 电气工程与系统科学 2025-07-25 Dou Hoon Kwark , Shirui Luo , Xiyue Zhu , Yudu Li , Zhi-Pei Liang , Volodymyr Kindratenko

Generative models have advanced significantly in realistic image synthesis, with diffusion models excelling in quality and stability. Recent multi-view diffusion models improve 3D-aware street view generation, but they struggle to produce…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Ji Li , Zhiwei Li , Shihao Li , Zhenjiang Yu , Boyang Wang , Haiou Liu

Recent advances in diffusion models have set an impressive milestone in many generation tasks, and trending works such as DALL-E2, Imagen, and Stable Diffusion have attracted great interest. Despite the rapid landscape changes, recent new…

计算机视觉与模式识别 · 计算机科学 2024-01-15 Xingqian Xu , Zhangyang Wang , Eric Zhang , Kai Wang , Humphrey Shi

Diffusion models have recently achieved remarkable performance in image super-resolution (SR), but their high computational cost limits practical deployment in remote sensing applications. To address this issue, we propose SlimDiffSR, a…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Ce Wang , Zhenyu Hu , Wanjie Sun

Accurate spatiotemporal pattern analysis is critical in fields such as urban traffic, meteorology, and public health monitoring. However, existing methods face performance bottlenecks, typically yielding only incremental gains and often…

机器学习 · 计算机科学 2026-05-20 Jing Chen , Shixiang Pan , Yujie Fan , Haocheng Ye , Haitao Xu , Wenqiang Xu

Deep learning-based image generation has undergone a paradigm shift since 2021, marked by fundamental architectural breakthroughs and computational innovations. Through reviewing architectural innovations and empirical results, this paper…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Benji Peng , Chia Xin Liang , Ziqian Bi , Ming Liu , Yichao Zhang , Tianyang Wang , Keyu Chen , Xinyuan Song , Pohsun Feng

In this paper, we present a new approach for model acceleration by exploiting spatial sparsity in visual data. We observe that the final prediction in vision Transformers is only based on a subset of the most informative tokens, which is…

计算机视觉与模式识别 · 计算机科学 2023-06-05 Yongming Rao , Zuyan Liu , Wenliang Zhao , Jie Zhou , Jiwen Lu

Transferring visual style between images while preserving semantic correspondence between similar objects remains a central challenge in computer vision. While existing methods have made great strides, most of them operate at global level…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Wenbo Nie , Zixiang Li , Renshuai Tao , Bin Wu , Yunchao Wei , Yao Zhao

Although remarkable progress has been made in image style transfer, style is just one of the components of artistic paintings. Directly transferring extracted style features to natural images often results in outputs with obvious synthetic…

计算机视觉与模式识别 · 计算机科学 2025-05-16 Nisha Huang , Weiming Dong , Yuxin Zhang , Fan Tang , Ronghui Li , Chongyang Ma , Xiu Li , Tong-Yee Lee , Changsheng Xu

Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial intelligence gap,''…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Shaoxiong Zhan , Yanlin Lai , Zheng Liu , Hai Lin , Shen Li , Xiaodong Cai , Zijian Lin , Wen Huang , Hai-Tao Zheng

In this paper, we propose SRIF, a novel Semantic shape Registration framework based on diffusion-based Image morphing and Flow estimation. More concretely, given a pair of extrinsically aligned shapes, we first render them from multi-views,…

计算机视觉与模式识别 · 计算机科学 2024-10-04 Mingze Sun , Chen Guo , Puhua Jiang , Shiwei Mao , Yurun Chen , Ruqi Huang

Diffusion models, known for their powerful generative capabilities, play a crucial role in addressing real-world super-resolution challenges. However, these models often focus on improving local textures while neglecting the impacts of…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Chunyang Bi , Xin Luo , Sheng Shen , Mengxi Zhang , Huanjing Yue , Jingyu Yang

For recent diffusion-based generative models, maintaining consistent content across a series of generated images, especially those containing subjects and complex details, presents a significant challenge. In this paper, we propose a new…

计算机视觉与模式识别 · 计算机科学 2024-05-03 Yupeng Zhou , Daquan Zhou , Ming-Ming Cheng , Jiashi Feng , Qibin Hou

The scarcity of annotated surgical data poses a significant challenge for developing deep learning systems in computer-assisted interventions. While diffusion models can synthesize realistic images, they often suffer from data memorization,…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Danush Kumar Venkatesh , Stefanie Speidel