中文
相关论文

相关论文: HetScene: Heterogeneity-Aware Diffusion for Dense …

200 篇论文

We explore different design choices for injecting noise into generative adversarial networks (GANs) with the goal of disentangling the latent space. Instead of traditional approaches, we propose feeding multiple noise codes through separate…

计算机视觉与模式识别 · 计算机科学 2020-05-06 Yazeed Alharbi , Peter Wonka

We study the problem of synthesizing immersive 3D indoor scenes from one or more images. Our aim is to generate high-resolution images and videos from novel viewpoints, including viewpoints that extrapolate far beyond the input images while…

计算机视觉与模式识别 · 计算机科学 2022-12-02 Jing Yu Koh , Harsh Agrawal , Dhruv Batra , Richard Tucker , Austin Waters , Honglak Lee , Yinfei Yang , Jason Baldridge , Peter Anderson

Compositional 3D scene synthesis has diverse applications across a spectrum of industries such as robotics, films, and video games, as it closely mirrors the complexity of real-world multi-object environments. Conventional works typically…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Yao Wei , Martin Renqiang Min , George Vosselman , Li Erran Li , Michael Ying Yang

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts for reasoning. However, the quality of existing interleaved…

We focus on explicitly learning disentangled representation for natural image generation, where the underlying spatial structure and the rendering on the structure can be independently controlled respectively, yet using no tuple…

机器学习 · 计算机科学 2019-10-01 Guang-Yuan Hao , Hong-Xing Yu , Wei-Shi Zheng

Reasoning about complex visual scenes involves perception of entities and their relations. Scene graphs provide a natural representation for reasoning tasks, by assigning labels to both entities (nodes) and relations (edges). Unfortunately,…

计算机视觉与模式识别 · 计算机科学 2020-03-17 Moshiko Raboh , Roei Herzig , Gal Chechik , Jonathan Berant , Amir Globerson

Imitation learning from human demonstrations is an effective paradigm for robot manipulation, but acquiring large datasets is costly and resource-intensive, especially for long-horizon tasks. To address this issue, we propose SkillMimicGen…

机器人学 · 计算机科学 2024-10-25 Caelan Garrett , Ajay Mandlekar , Bowen Wen , Dieter Fox

Addressing the challenge of removing atmospheric fog or haze from digital images, known as image dehazing, has recently gained significant traction in the computer vision community. Although contemporary dehazing models have demonstrated…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Anas M. Ali , Anis Koubaa , Bilel Benjdira

Dense indoor scene modeling from 2D images has been bottlenecked due to the absence of depth information and cluttered occlusions. We present an automatic indoor scene modeling approach using deep features from neural networks. Given a…

计算机视觉与模式识别 · 计算机科学 2020-02-25 Yinyu Nie , Shihui Guo , Jian Chang , Xiaoguang Han , Jiahui Huang , Shi-Min Hu , Jian Jun Zhang

Human vision combines low-resolution "gist" information from the visual periphery with sparse but high-resolution information from fixated locations to construct a coherent understanding of a visual scene. In this paper, we introduce…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Ritik Raina , Abe Leite , Alexandros Graikos , Seoyoung Ahn , Dimitris Samaras , Gregory J. Zelinsky

Accurately reconstructing complex full multi-object scenes from sparse observations remains a core challenge in computer vision and a key step toward scalable and reliable simulation for robotics. In this work, we introduce RecGen, a…

Learning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characteristics of the existing datasets on Human-Scene Interaction (HSI); they only have limited scale/quality…

计算机视觉与模式识别 · 计算机科学 2022-10-19 Zan Wang , Yixin Chen , Tengyu Liu , Yixin Zhu , Wei Liang , Siyuan Huang

We seek to give users precise control over diffusion-based image generation by modeling complex scenes as sequences of layers, which define the desired spatial arrangement and visual attributes of objects in the scene. Collage Diffusion…

计算机视觉与模式识别 · 计算机科学 2023-09-01 Vishnu Sarukkai , Linden Li , Arden Ma , Christopher Ré , Kayvon Fatahalian

Remote sensing vision tasks require extensive labeled data across multiple, interconnected domains. However, current generative data augmentation frameworks are task-isolated, i.e., each vision task requires training an independent…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Datao Tang , Hao Wang , Yudeng Xin , Hui Qiao , Dongsheng Jiang , Yin Li , Zhiheng Yu , Xiangyong Cao

As smart homes become increasingly prevalent, intelligent models are widely used for tasks such as anomaly detection and behavior prediction. These models are typically trained on static datasets, making them brittle to behavioral drift…

人工智能 · 计算机科学 2025-08-06 Zhiyao Xu , Dan Zhao , Qingsong Zou , Qing Li , Yong Jiang , Yuhang Wang , Jingyu Xiao

Three-dimensional building generation is vital for applications in gaming, virtual reality, and digital twins, yet current methods face challenges in producing diverse, structured, and hierarchically coherent buildings. We propose…

图形学 · 计算机科学 2025-05-08 Junming Huang , Chi Wang , Letian Li , Changxin Huang , Qiang Dai , Weiwei Xu

Controllable layout generation aims at synthesizing plausible arrangement of element bounding boxes with optional constraints, such as type or position of a specific element. In this work, we try to solve a broad range of layout generation…

计算机视觉与模式识别 · 计算机科学 2023-03-15 Naoto Inoue , Kotaro Kikuchi , Edgar Simo-Serra , Mayu Otani , Kota Yamaguchi

We study to generate novel views of indoor scenes given sparse input views. The challenge is to achieve both photorealism and view consistency. We present SparseGNV: a learning framework that incorporates 3D structures and image generative…

计算机视觉与模式识别 · 计算机科学 2023-05-12 Weihao Cheng , Yan-Pei Cao , Ying Shan

Long-context training of large language models (LLMs) is commonly distributed with Context Parallelism (CP) and Head Parallelism (HP), but existing training systems largely assume homogeneous GPU meshes. This paper extends CP and HP to…

分布式、并行与集群计算 · 计算机科学 2026-05-11 Yan Liang , Youhe Jiang , Ran Yan , Binhang Yuan , Wei Wang , Chuan Wu

Multi-subject image generation requires seamlessly harmonizing multiple reference identities within a coherent scene. However, existing methods relying on rigid spatial masks or localized attention often struggle with the…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Honghao Cai , Xiangyuan Wang , Jing Li , Yunhao Bai , Tianze Zhou , Haohua Chen , Chao Hui , Changhao Qiao , Runqi Wang , Sijie Xu , Yuyang Hao , Zezhou Cui , Yuyuan Yang , Wei Zhu , Yibo Chen , Xu Tang , Yao Hu , Zhen Li