中文
相关论文

相关论文: ATISS: Autoregressive Transformers for Indoor Scen…

200 篇论文

A compositional understanding of the world in terms of objects and their geometry in 3D space is considered a cornerstone of human cognition. Facilitating the learning of such a representation in neural networks holds promise for…

Diffusion models have demonstrated impressive capabilities in modeling complex data distributions and are increasingly applied in various generative tasks. In this work, we propose Pose Analysis by Diffusion Synthesis PADS, a unified…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Haorui Ji , Hongdong Li

Detecting a diverse range of objects under various driving scenarios is essential for the effectiveness of autonomous driving systems. However, the real-world data collected often lacks the necessary diversity presenting a long-tail…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Aqeel Anwar , Tae Eun Choe , Zian Wang , Sanja Fidler , Minwoo Park

Indoor scene recognition is a multi-faceted and challenging problem due to the diverse intra-class variations and the confusing inter-class similarities. This paper presents a novel approach which exploits rich mid-level convolutional…

计算机视觉与模式识别 · 计算机科学 2016-06-29 Salman H. Khan , Munawar Hayat , Mohammed Bennamoun , Roberto Togneri , Ferdous Sohel

The volume and diversity of training data are critical for modern deep learningbased methods. Compared to the massive amount of labeled perspective images, 360 panoramic images fall short in both volume and diversity. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2023-09-28 Yu-Cheng Hsieh , Cheng Sun , Suraj Dengale , Min Sun

Recent advances in implicit scene representation enable high-fidelity street view novel view synthesis. However, existing methods optimize a neural radiance field for each scene, relying heavily on dense training images and extensive…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Sheng Miao , Jiaxin Huang , Dongfeng Bai , Weichao Qiu , Bingbing Liu , Andreas Geiger , Yiyi Liao

Furnishing and rendering indoor scenes has been a long-standing task for interior design, where artists create a conceptual design for the space, build a 3D model of the space, decorate, and then perform rendering. Although the task is…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Hong-Wing Pang , Yingshu Chen , Phuoc-Hieu Le , Binh-Son Hua , Duc Thanh Nguyen , Sai-Kit Yeung

Image generation models trained on large datasets can synthesize high-quality images but often produce spatially inconsistent and distorted images due to limited information about the underlying structures and spatial layouts. In this work,…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Hyundo Lee , Suhyung Choi , Inwoo Hwang , Byoung-Tak Zhang

Modeling the dynamic behavior of deformable objects is crucial for creating realistic digital worlds. While conventional simulations produce high-quality motions, their computational costs are often prohibitive. Subspace simulation…

We introduce the GANformer, a novel and efficient type of transformer, and explore it for the task of visual generative modeling. The network employs a bipartite structure that enables long-range interactions across the image, while…

计算机视觉与模式识别 · 计算机科学 2022-03-30 Drew A. Hudson , C. Lawrence Zitnick

Simulating stylized human-scene interactions (HSI) in physical environments is a challenging yet fascinating task. Prior works emphasize long-term execution but fall short in achieving both diverse style and physical plausibility. To tackle…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Wenjia Wang , Liang Pan , Zhiyang Dou , Jidong Mei , Zhouyingcheng Liao , Yuke Lou , Yifan Wu , Lei Yang , Jingbo Wang , Taku Komura

Transformers, which were originally developed for natural language processing, have recently generated significant interest in the computer vision and audio communities due to their flexibility in learning long-range relationships.…

声音 · 计算机科学 2024-08-15 Sara Atito , Muhammad Awais , Wenwu Wang , Mark D Plumbley , Josef Kittler

We present a novel task, i.e., animating a target 3D object through the motion of a raw driving sequence. In previous works, extra auxiliary correlations between source and target meshes or intermedia factors are inevitable to capture the…

计算机视觉与模式识别 · 计算机科学 2021-10-22 Haoyu Chen , Hao Tang , Nicu Sebe , Guoying Zhao

This paper presents a novel approach to neural instrument sound synthesis using a two-stage semi-supervised learning framework capable of generating pitch-accurate, high-quality music samples from an expressive timbre latent space. Existing…

声音 · 计算机科学 2025-10-07 Christian Limberg , Fares Schulz , Zhe Zhang , Stefan Weinzierl

Although recently several foundation models for satellite remote sensing imagery have been proposed, they fail to address major challenges of real/operational applications. Indeed, embeddings that don't take into account the spectral,…

人工智能 · 计算机科学 2024-10-01 Iris Dumeur , Silvia Valero , Jordi Inglada

We describe a novel learning-by-synthesis method for estimating gaze direction of an automated intelligent surveillance system. Recently, progress in learning-by-synthesis has proposed training models on synthetic images, which can…

计算机视觉与模式识别 · 计算机科学 2018-10-09 Tongtong Zhao , Yuxiao Yan , Jinjia Peng , Zetian Mi , Xianping Fu

In this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimodal framework capable of handling the speech and text…

This paper presents ARTEMIS, an end-to-end autonomous driving framework that combines autoregressive trajectory planning with Mixture-of-Experts (MoE). Traditional modular methods suffer from error propagation, while existing end-to-end…

机器人学 · 计算机科学 2025-05-06 Renju Feng , Ning Xi , Duanfeng Chu , Rukang Wang , Zejian Deng , Anzheng Wang , Liping Lu , Jinxiang Wang , Yanjun Huang

Neural rendering techniques promise efficient photo-realistic image synthesis while at the same time providing rich control over scene parameters by learning the physical image formation process. While several supervised methods have been…

计算机视觉与模式识别 · 计算机科学 2021-03-30 Hassan Abu Alhaija , Siva Karthik Mustikovela , Justus Thies , Varun Jampani , Matthias Nießner , Andreas Geiger , Carsten Rother

Designing 3D scenes is traditionally a challenging task that demands both artistic expertise and proficiency with complex software. Recent advances in text-to-3D generation have greatly simplified this process by letting users create scenes…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Zeqi Gu , Yin Cui , Zhaoshuo Li , Fangyin Wei , Yunhao Ge , Jinwei Gu , Ming-Yu Liu , Abe Davis , Yifan Ding