English
Related papers

Related papers: 3D-TOGO: Towards Text-Guided Cross-Category 3D Obj…

200 papers

This paper presents Matrix, an advanced AI-powered framework designed for real-time 3D object generation in Augmented Reality (AR) environments. By integrating a cutting-edge text-to-3D generative AI model, multilingual speech-to-text…

Human-Computer Interaction · Computer Science 2025-03-24 Majid Behravan , Denis Gracanin

Text-to-Image (T2I) models excel at synthesizing concepts such as nouns, appearances, and styles. To enable customized content creation based on a few example images of a concept, methods such as Textual Inversion and DreamBooth invert the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-30 Saman Motamed , Danda Pani Paudel , Luc Van Gool

Current audio-driven 3D head generation methods mainly focus on single-speaker scenarios, lacking natural, bidirectional listen-and-speak interaction. Achieving seamless conversational behavior, where speaking and listening states…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Lei Zhu , Lijian Lin , Ye Zhu , Jiahao Wu , Xuehan Hou , Yu Li , Yunfei Liu , Jie Chen

We investigate the incorporation of visual relationships into the task of supervised image caption generation by proposing a model that leverages detected objects and auto-generated visual relationships to describe images in natural…

Computer Vision and Pattern Recognition · Computer Science 2021-09-24 Maximilian Mozes , Martin Schmitt , Vladimir Golkov , Hinrich Schütze , Daniel Cremers

Recent approaches have achieved great success in image generation from structured inputs, e.g., semantic segmentation, scene graph or layout. Although these methods allow specification of objects and their locations at image-level, they…

Computer Vision and Pattern Recognition · Computer Science 2020-08-28 Ke Ma , Bo Zhao , Leonid Sigal

For humans, visual understanding is inherently generative: given a 3D shape, we can postulate how it would look in the world; given a 2D image, we can infer the 3D structure that likely gave rise to it. We can thus translate between the 2D…

Computer Vision and Pattern Recognition · Computer Science 2020-11-17 Tristan Aumentado-Armstrong , Alex Levinshtein , Stavros Tsogkas , Konstantinos G. Derpanis , Allan D. Jepson

Achieving visual semantic understanding requires a unified framework that simultaneously handles object detection, category prediction, and attribute recognition. However, current advanced approaches rely on global similarity and struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Xinyu Nan , Lingtao Mao , Huangyu Dai , Zexin Zheng , Xinyu Sun , Zihan Liang , Ben Chen , Yuqing Ding , Chenyi Lei , Wenwu Ou , Han Li

Text-to-3D generation has recently seen significant progress. To enhance its practicality in real-world applications, it is crucial to generate multiple independent objects with interactions, similar to layer-compositing in 2D image…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Zizheng Yan , Jiapeng Zhou , Fanpeng Meng , Yushuang Wu , Lingteng Qiu , Zisheng Ye , Shuguang Cui , Guanying Chen , Xiaoguang Han

Despite the unprecedented success of text-to-image diffusion models, controlling the number of depicted objects using text is surprisingly hard. This is important for various applications from technical documents, to children's books to…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Lital Binyamin , Yoad Tewel , Hilit Segev , Eran Hirsch , Royi Rassin , Gal Chechik

The entertainment industry relies on 3D visual content to create immersive experiences, but traditional methods for creating textured 3D models can be time-consuming and subjective. Generative networks such as StyleGAN have advanced image…

Computer Vision and Pattern Recognition · Computer Science 2024-02-09 Yi-Ting Pan , Chai-Rong Lee , Shu-Ho Fan , Jheng-Wei Su , Jia-Bin Huang , Yung-Yu Chuang , Hung-Kuo Chu

Current state-of-the-art methods for text-to-shape generation either require supervised training using a labeled dataset of pre-defined 3D shapes, or perform expensive inference-time optimization of implicit neural representations. In this…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Kelly O. Marshall , Minh Pham , Ameya Joshi , Anushrut Jignasu , Aditya Balu , Adarsh Krishnamurthy , Chinmay Hegde

We introduce InseRF, a novel method for generative object insertion in the NeRF reconstructions of 3D scenes. Based on a user-provided textual description and a 2D bounding box in a reference viewpoint, InseRF generates new objects in 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-01-11 Mohamad Shahbazi , Liesbeth Claessens , Michael Niemeyer , Edo Collins , Alessio Tonioni , Luc Van Gool , Federico Tombari

Diffusion models have demonstrated impressive performance in text-to-image generation. They utilize a text encoder and cross-attention blocks to infuse textual information into images at a pixel level. However, their capability to generate…

Computer Vision and Pattern Recognition · Computer Science 2023-06-06 Luping Liu , Zijian Zhang , Yi Ren , Rongjie Huang , Xiang Yin , Zhou Zhao

This paper presents a Generative RegIon-to-Text transformer, GRiT, for object understanding. The spirit of GRiT is to formulate object understanding as <region, text> pairs, where region locates objects and text describes objects. For…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Jialian Wu , Jianfeng Wang , Zhengyuan Yang , Zhe Gan , Zicheng Liu , Junsong Yuan , Lijuan Wang

Texture editing is a crucial task in 3D modeling that allows users to automatically manipulate the surface materials of 3D models. However, the inherent complexity of 3D models and the ambiguous text description lead to the challenge in…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Shengqi Liu , Zhuo Chen , Jingnan Gao , Yichao Yan , Wenhan Zhu , Jiangjing Lyu , Xiaokang Yang

Text-to-3D generation is to craft a 3D object according to a natural language description. This can significantly reduce the workload for manually designing 3D models and provide a more natural way of interaction for users. However, this…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Han Yi , Zhedong Zheng , Xiangyu Xu , Tat-seng Chua

We identify occlusion reasoning as a fundamental yet overlooked aspect for 3D layout-conditioned generation. It is essential for synthesizing partially occluded objects with depth-consistent geometry and scale. While existing methods can…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Vaibhav Agrawal , Rishubh Parihar , Pradhaan Bhat , Ravi Kiran Sarvadevabhatla , R. Venkatesh Babu

Text-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the target objects and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Xingrui Wang , Xin Li , Yaosi Hu , Hanxin Zhu , Chen Hou , Cuiling Lan , Zhibo Chen

Image caption generation is one of the most challenging problems at the intersection of vision and language domains. In this work, we propose a realistic captioning task where the input scenes may incorporate visual objects with no…

Computer Vision and Pattern Recognition · Computer Science 2022-07-04 Berkan Demirel , Ramazan Gokberk Cinbis

Many high-level video understanding methods require input in the form of object proposals. Currently, such proposals are predominantly generated with the help of networks that were trained for detecting and segmenting a set of known object…

Computer Vision and Pattern Recognition · Computer Science 2020-05-22 Aljosa Osep , Paul Voigtlaender , Mark Weber , Jonathon Luiten , Bastian Leibe
‹ Prev 1 4 5 6 7 8 10 Next ›