English
Related papers

Related papers: MMGDreamer: Mixed-Modality Graph for Geometry-Cont…

200 papers

Diffusion-based methods have achieved prominent success in generating 2D media. However, accomplishing similar proficiencies for scene-level mesh texturing in 3D spatial applications, e.g., XR/VR, remains constrained, primarily due to the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-23 Bangbang Yang , Wenqi Dong , Lin Ma , Wenbo Hu , Xiao Liu , Zhaopeng Cui , Yuewen Ma

A Scene, represented visually using different formats such as RGB-D, LiDAR scan, keypoints, rectangular, spherical, multi-views, etc., contains information implicitly embedded relevant to applications such as scene indexing, vision-based…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Preeti Meena , Himanshu Kumar , Sandeep Yadav

LiDAR and camera are two modalities available for 3D semantic segmentation in autonomous driving. The popular LiDAR-only methods severely suffer from inferior segmentation on small and distant objects due to insufficient laser points, while…

Computer Vision and Pattern Recognition · Computer Science 2023-03-16 Jiale Li , Hang Dai , Hao Han , Yong Ding

Diffusion models are advancing autonomous driving by enabling realistic data synthesis, predictive end-to-end planning, and closed-loop simulation, with a primary focus on temporally consistent generation. However, large-scale 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Yu Yang , Alan Liang , Jianbiao Mei , Yukai Ma , Yong Liu , Gim Hee Lee

Mobile manipulators in households must both navigate and manipulate. This requires a compact, semantically rich scene representation that captures where objects are, how they function, and which parts are actionable. Scene graphs are a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Yuanchen Ju , Yongyuan Liang , Yen-Jen Wang , Nandiraju Gireesh , Yuanliang Ju , Seungjae Lee , Qiao Gu , Elvis Hsieh , Furong Huang , Koushil Sreenath

Embodied navigation is a fundamental capability for robotic agents operating. Real-world deployment requires open vocabulary generalization and low training overhead, motivating zero-shot methods rather than task-specific RL training.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Xun Huang , Shijia Zhao , Yunxiang Wang , Xin Lu , Wanfa Zhang , Rongsheng Qu , Weixin Li , Yunhong Wang , Chenglu Wen

Multimodal learning combines multiple data modalities, broadening the types and complexity of data our models can utilize: for example, from plain text to image-caption pairs. Most multimodal learning algorithms focus on modeling simple…

Artificial Intelligence · Computer Science 2023-10-13 Minji Yoon , Jing Yu Koh , Bryan Hooi , Ruslan Salakhutdinov

Recent advances in Diffusion Probabilistic Models (DPMs) have set new standards in high-quality image synthesis. Yet, controlled generation remains challenging, particularly in sensitive areas such as medical imaging. Medical images feature…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Sarah Cechnicka , Matthew Baugh , Weitong Zhang , Mischa Dombrowski , Zhe Li , Johannes C. Paetzold , Candice Roufosse , Bernhard Kainz

Scene graphs offer a structured, hierarchical representation of images, with nodes and edges symbolizing objects and the relationships among them. It can serve as a natural interface for image editing, dramatically improving precision and…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Zhiyuan Zhang , DongDong Chen , Jing Liao

Recent multimodal face generation models address the spatial control limitations of text-to-image diffusion models by augmenting text-based conditioning with spatial priors such as segmentation masks, sketches, or edge maps. This multimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Bharath Krishnamurthy , Ajita Rattani

There has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple objects and their…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Yunnan Wang , Ziqiang Li , Zequn Zhang , Wenyao Zhang , Baao Xie , Xihui Liu , Wenjun Zeng , Xin Jin

Neural fields have achieved impressive advancements in view synthesis and scene reconstruction. However, editing these neural fields remains challenging due to the implicit encoding of geometry and texture information. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2023-09-08 Jingyu Zhuang , Chen Wang , Lingjie Liu , Liang Lin , Guanbin Li

Multi-view image generation holds significant application value in computer vision, particularly in domains like 3D reconstruction, virtual reality, and augmented reality. Most existing methods, which rely on extending single images, face…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Jiaqi Wu , Yaosen Chen , Shuyuan Zhu

Traditional 3D garment creation is labor-intensive, involving sketching, modeling, UV mapping, and texturing, which are time-consuming and costly. Recent advances in diffusion-based generative models have enabled new possibilities for 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Boqian Li , Xuan Li , Ying Jiang , Tianyi Xie , Feng Gao , Huamin Wang , Yin Yang , Chenfanfu Jiang

Implicit neural representation has demonstrated promising results in 3D reconstruction on various scenes. However, existing approaches either struggle to model fast-moving objects or are incapable of handling large-scale camera ego-motions…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Tianchen Deng , Yanbo Wang , Yejia Liu , Chenpeng Su , Jingchuan Wang , Danwei Wang , Shao-Yuan Lo , Weidong Chen

We propose a novel task of text-controlled human object interaction generation in 3D scenes with movable objects. Existing human-scene interaction datasets suffer from insufficient interaction categories and typically only consider…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Xinhao Cai , Minghang Zheng , Xin Jin , Yang Liu

Large-scale video generative models can synthesize diverse and realistic visual content for dynamic world creation, but they often lack element-wise controllability, hindering their use in editing scenes and training embodied AI agents. We…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Sicheng Mo , Ziyang Leng , Leon Liu , Weizhen Wang , Honglin He , Bolei Zhou

Advances in image diffusion models have recently led to notable improvements in the generation of high-quality images. In combination with Neural Radiance Fields (NeRFs), they enabled new opportunities in 3D generation. However, most…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Jan-Niklas Dihlmann , Andreas Engelhardt , Hendrik Lensch

Scene Graph Generation (SGG) serves a comprehensive representation of the images for human understanding as well as visual understanding tasks. Due to the long tail bias problem of the object and predicate labels in the available annotated…

Computer Vision and Pattern Recognition · Computer Science 2022-11-10 Anh Duc Bui , Soyeon Caren Han , Josiah Poon

Benefiting from the powerful expressive capability of graphs, graph-based approaches have achieved impressive performance in various biomedical applications. Most existing methods tend to define the adjacency matrix among samples manually…

Machine Learning · Computer Science 2021-07-02 Shuai Zheng , Zhenfeng Zhu , Zhizhe Liu , Zhenyu Guo , Yang Liu , Yao Zhao
‹ Prev 1 4 5 6 7 8 10 Next ›