English
Related papers

Related papers: Layout Anything: One Transformer for Universal Roo…

200 papers

This work is motivated by recent developments in Deep Neural Networks, particularly the Transformer architectures underlying applications such as ChatGPT, and the need for performing inference on mobile devices. Focusing on emerging…

Machine Learning · Computer Science 2024-04-23 Wei Niu , Md Musfiqur Rahman Sanim , Zhihao Shu , Jiexiong Guan , Xipeng Shen , Miao Yin , Gagan Agrawal , Bin Ren

We present a general-purpose framework for image modelling and vision tasks based on probabilistic frame prediction. Our approach unifies a broad range of tasks, from image segmentation, to novel view synthesis and video interpolation. We…

Computer Vision and Pattern Recognition · Computer Science 2022-05-10 Charlie Nash , João Carreira , Jacob Walker , Iain Barr , Andrew Jaegle , Mateusz Malinowski , Peter Battaglia

Existing unified image segmentation models either employ a unified architecture across multiple tasks but use separate weights tailored to each dataset, or apply a single set of weights to multiple datasets but are limited to a single task.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Pei Wang , Zhaowei Cai , Hao Yang , Ashwin Swaminathan , R. Manmatha , Stefano Soatto

In remote sensing there exists a common need for learning scale invariant shapes of objects like buildings. Prior works relies on tweaking multiple loss functions to convert segmentation maps into the final scale invariant representation,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Maxim Khomiakov , Michael Riis Andersen , Jes Frellsen

This paper presents a neural network built upon Transformers, namely PlaneTR, to simultaneously detect and reconstruct planes from a single image. Different from previous methods, PlaneTR jointly leverages the context information and the…

Computer Vision and Pattern Recognition · Computer Science 2021-07-29 Bin Tan , Nan Xue , Song Bai , Tianfu Wu , Gui-Song Xia

Depth completion aims to predict dense depth maps with sparse depth measurements from a depth sensor. Currently, Convolutional Neural Network (CNN) based models are the most popular methods applied to depth completion tasks. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Jian Qian , Miao Sun , Ashley Lee , Jie Li , Shenglong Zhuo , Patrick Yin Chiang

Real-world exposure correction is fundamentally challenged by spatially non-uniform degradations, where diverse exposure errors frequently coexist within a single image. However, existing exposure correction methods are still largely…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Ao Li , Jiawei Sun , Le Dong , Zhenyu Wang , Weisheng Dong

We address the problem of scene layout generation for diverse domains such as images, mobile applications, documents, and 3D objects. Most complex scenes, natural or human-designed, can be expressed as a meaningful arrangement of simpler…

Computer Vision and Pattern Recognition · Computer Science 2021-10-01 Kamal Gupta , Justin Lazarow , Alessandro Achille , Larry Davis , Vijay Mahadevan , Abhinav Shrivastava

Current 6D object pose estimation methods usually require a 3D model for each object. These methods also require additional training in order to incorporate new objects. As a result, they are difficult to scale to a large number of objects…

Computer Vision and Pattern Recognition · Computer Science 2020-06-15 Keunhong Park , Arsalan Mousavian , Yu Xiang , Dieter Fox

A unified simulator that can model diverse physical phenomena without solver-specific redesign is a long-standing goal across simulation science. We present a learning-based particle simulator built on a single transformer architecture to…

Modern scene reconstruction methods are able to accurately recover 3D surfaces that are visible in one or more images. However, this leads to incomplete reconstructions, missing all occluded surfaces. While much progress has been made on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Sam Bahrami , Dylan Campbell

In this work, we aim to improve the 3D reasoning ability of Transformers in multi-view 3D human pose estimation. Recent works have focused on end-to-end learning-based transformer designs, which struggle to resolve geometric information…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Ziwei Liao , Jialiang Zhu , Chunyu Wang , Han Hu , Steven L. Waslander

In this paper, we present Uformer, an effective and efficient Transformer-based architecture for image restoration, in which we build a hierarchical encoder-decoder network using the Transformer block. In Uformer, there are two core…

Computer Vision and Pattern Recognition · Computer Science 2021-11-29 Zhendong Wang , Xiaodong Cun , Jianmin Bao , Wengang Zhou , Jianzhuang Liu , Houqiang Li

We propose a method for room layout estimation that does not rely on the typical box approximation or Manhattan world assumption. Instead, we reformulate the geometry inference problem as an instance detection task, which we solve by…

Computer Vision and Pattern Recognition · Computer Science 2019-05-09 Henry Howard-Jenkins , Shuda Li , Victor Prisacariu

Current 3D layout estimation models are primarily trained on synthetic datasets containing simple single room or single floor environments. As a consequence, they cannot natively handle large multi floor buildings and require scenes to be…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Valentin Bieri , Marie-Julie Rakotosaona , Keisuke Tateno , Francis Engelmann , Leonidas Guibas

We investigate the problem of automatically placing an object into a background image for image compositing. Given a background image and a segmented object, the goal is to train a model to predict plausible placements (location and scale)…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Sijie Zhu , Zhe Lin , Scott Cohen , Jason Kuen , Zhifei Zhang , Chen Chen

We propose LGGPT, an LLM-based model tailored for unified layout generation. First, we propose Arbitrary Layout Instruction (ALI) and Universal Layout Response (ULR) as the uniform I/O template. ALI accommodates arbitrary layout generation…

Machine Learning · Computer Science 2025-02-21 Peirong Zhang , Jiaxin Zhang , Jiahuan Cao , Hongliang Li , Lianwen Jin

Unlike other vision tasks where Transformer-based approaches are becoming increasingly common, stereo depth estimation is still dominated by convolution-based approaches. This is mainly due to the limited availability of real-world ground…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Soomin Kim , Hyesong Choi , Jihye Ahn , Dongbo Min

We introduce, design, and evaluate a set of universal receiver beamforming techniques. Our approach and system DEFORM, a Deep Learning (DL) based RX beamforming achieves significant gain for multi antenna RF receivers while being agnostic…

Networking and Internet Architecture · Computer Science 2022-03-21 Hai N. Nguyen , Guevara Noubir

This paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is -- learning diffusion models for marginal, conditional, and joint…

Machine Learning · Computer Science 2023-05-31 Fan Bao , Shen Nie , Kaiwen Xue , Chongxuan Li , Shi Pu , Yaole Wang , Gang Yue , Yue Cao , Hang Su , Jun Zhu
‹ Prev 1 3 4 5 6 7 10 Next ›