English
Related papers

Related papers: Beyond Spatial Compression: Interface-Centric Gene…

200 papers

In autoregressive (AR) image generation, visual tokenizers compress images into compact discrete latent tokens, enabling efficient training of downstream autoregressive models for visual generation via next-token prediction. While scaling…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Tianwei Xiong , Jun Hao Liew , Zilong Huang , Jiashi Feng , Xihui Liu

Object-centric representations promise a key property for few-shot learning: Rather than treating a scene as a single unit, a model can decompose it into individual object-level parts that can be matched and compared across different…

Machine Learning · Computer Science 2026-05-19 Phu-Quy Nguyen-Lam , Phu-Hoa Pham , Dao Sy Duy Minh , Chi-Nguyen Tran , Huynh Trung Kiet , Long Tran-Thanh

We introduce a novel 3D generation method for versatile and high-quality 3D asset creation. The cornerstone is a unified Structured LATent (SLAT) representation which allows decoding to different output formats, such as Radiance Fields, 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Jianfeng Xiang , Zelong Lv , Sicheng Xu , Yu Deng , Ruicheng Wang , Bowen Zhang , Dong Chen , Xin Tong , Jiaolong Yang

Composition-the ability to generate myriad variations from finite means-is believed to underlie powerful generalization. However, compositional generalization remains a key challenge for deep learning. A widely held assumption is that…

Machine Learning · Computer Science 2025-05-27 Qiyao Liang , Daoyuan Qian , Liu Ziyin , Ila Fiete

In the majority of GAN architectures, the latent space is defined as a set of vectors of given dimensionality. Such representations are not easily interpretable and do not capture spatial information of image content directly. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Maciej Sypetkowski

Diffusion-based generative image compression has demonstrated remarkable potential for achieving realistic reconstruction at ultra-low bitrates. The key to unlocking this potential lies in making the entire compression process…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Xihua Sheng , Lingyu Zhu , Tianyu Zhang , Dong Liu , Shiqi Wang , Jing Wang

Current 3D-aware pretraining methods for embodied perception and manipulation are largely built on differentiable rendering frameworks, producing either fully implicit neural fields or fully explicit geometric primitives. Implicit…

We propose a 3D face generative model with local weights to increase the model's variations and expressiveness. The proposed model allows partial manipulation of the face while still learning the whole face mesh. For this purpose, we…

Graphics · Computer Science 2021-07-20 Minyoung Kim , Young J. Kim

Linear attention has emerged as a promising direction for scaling Vision Transformers beyond the quadratic cost of dense self-attention. A prevalent strategy is to compress spatial tokens into a compact set of intermediate proxies that…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yuntong Li , Hainuo Wang , Hengxing Liu , Mingjia Li , Xiaojie Guo

3D open-world classification is a challenging yet essential task in dynamic and unstructured real-world scenarios, requiring both open-category and open-pose recognition. To address these challenges, recent wisdom often takes sophisticated…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Xinzhe Xia , Weiguang Zhao , Yuyao Yan , Guanyu Yang , Rui Zhang , Kaizhu Huang , Xi Yang

Deep learning models have significantly improved the visual quality and accuracy on compressive sensing recovery. In this paper, we propose an algorithm for signal reconstruction from compressed measurements with image priors captured by a…

Machine Learning · Computer Science 2020-03-20 Shaojie Xu , Sihan Zeng , Justin Romberg

Recent AI-based 3D content creation has largely evolved along two paths: feed-forward image-to-3D reconstruction approaches and 3D generative models trained with 2D or 3D supervision. In this work, we show that existing feed-forward…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Suttisak Wizadwongsa , Jinfan Zhou , Edward Li , Jeong Joon Park

Understanding and generating 3D objects as compositions of meaningful parts is fundamental to human perception and reasoning. However, most text-to-3D methods overlook the semantic and functional structure of parts. While recent part-aware…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Tianjiao Yu , Xinzhuo Li , Muntasir Wahed , Jerry Xiong , Yifan Shen , Ying Shen , Ismini Lourentzou

3D generation has witnessed significant advancements, yet efficiently producing high-quality 3D assets from a single image remains challenging. In this paper, we present a triplane autoencoder, which encodes 3D models into a compact…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Bowen Zhang , Tianyu Yang , Yu Li , Lei Zhang , Xi Zhao

Photo-realistic and controllable 3D avatars are crucial for various applications such as virtual and mixed reality (VR/MR), telepresence, gaming, and film production. Traditional methods for avatar creation often involve time-consuming…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Keqiang Sun , Amin Jourabloo , Riddhish Bhalodia , Moustafa Meshry , Yu Rong , Zhengyu Yang , Thu Nguyen-Phuoc , Christian Haene , Jiu Xu , Sam Johnson , Hongsheng Li , Sofien Bouaziz

3D Scene Graph (3DSG) generation plays a pivotal role in spatial understanding and affordance perception. To mitigate generalization issues from data scarcity, joint-embedding and generative proxy tasks are proposed to pre-train 3DSG…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Yucheng Huang , Luping Ji , Xiangwei Jiang , Wen Li , Mao Ye

We propose TC-AE, a ViT-based architecture for deep compression autoencoders. Existing methods commonly increase the channel number of latent representations to maintain reconstruction quality under high compression ratios. However, this…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Teng Li , Ziyuan Huang , Cong Chen , Yangfu Li , Yuanhuiyi Lyu , Dandan Zheng , Chunhua Shen , Jun Zhang

Image denoising aims to remove noise while preserving structural details and perceptual realism, yet distortion-driven methods often produce over-smoothed reconstructions, especially under strong noise and distribution shift. This paper…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Nam Nguyen , Thinh Nguyen , Bella Bose

Encoder transformer models compress information from all tokens in a sequence into a single [CLS] token to represent global context. This approach risks diluting fine-grained or hierarchical features, leading to information loss in…

Computation and Language · Computer Science 2025-09-23 Asif Shahriar , Rifat Shahriyar , M Saifur Rahman

In order to simulate the mechanical behavior of large structures assembled from thin composite panels, we propose a coupling technique which substitutes local 3D models for the global plate model in the critical zones where plate modeling…

Numerical Analysis · Mathematics 2015-01-09 Guillaume Guguin , Olivier Allix , Pierre Gosselet , Stéphane Guinard