中文
相关论文

相关论文: 3D-TOGO: Towards Text-Guided Cross-Category 3D Obj…

200 篇论文

In this paper, we develop a new method, termed SDF-3DGAN, for 3D object generation and 3D-Aware image synthesis tasks, which introduce implicit Signed Distance Function (SDF) as the 3D object representation method in the generative field.…

计算机视觉与模式识别 · 计算机科学 2023-03-14 Lutao Jiang , Ruyi Ji , Libo Zhang

Current image generation and editing methods primarily process textual prompts as direct inputs without reasoning about visual composition and explicit operations. We present Generation Chain-of-Thought (GoT), a novel paradigm that enables…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Rongyao Fang , Chengqi Duan , Kun Wang , Linjiang Huang , Hao Li , Shilin Yan , Hao Tian , Xingyu Zeng , Rui Zhao , Jifeng Dai , Xihui Liu , Hongsheng Li

Reconstructing a renderable 3D model from images is a useful but challenging task. Recent feedforward 3D reconstruction methods have demonstrated remarkable success in efficiently recovering geometry, but still cannot accurately model the…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Zekai Gu , Shuoxuan Feng , Yansong Wang , Hanzhuo Huang , Zhongshuo Du , Chengfeng Zhao , Chengwei Ren , Peng Wang , Yuan Liu

Synthesizing multi-view 3D from one single image is a significant but challenging task. Zero-1-to-3 methods have achieved great success by lifting a 2D latent diffusion model to the 3D scope. The target view image is generated with a…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Yabo Chen , Jiemin Fang , Yuyang Huang , Taoran Yi , Xiaopeng Zhang , Lingxi Xie , Xinggang Wang , Wenrui Dai , Hongkai Xiong , Qi Tian

We propose a novel task of text-controlled human object interaction generation in 3D scenes with movable objects. Existing human-scene interaction datasets suffer from insufficient interaction categories and typically only consider…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Xinhao Cai , Minghang Zheng , Xin Jin , Yang Liu

The recently emerging conditional diffusion models seem promising for mitigating the labor and expenses in building large 3D medical imaging datasets. However, previous studies on 3D CT generation primarily focus on specific organs…

图像与视频处理 · 电气工程与系统科学 2025-12-02 Linrui Dai , Rongzhao Zhang , Yongrui Yu , Xiaofan Zhang

Learning to use tools or objects in common scenes, particularly handling them in various ways as instructed, is a key challenge for developing interactive robots. Training models to generate such manipulation trajectories requires a large…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

Text-to-image (TTI) diffusion models have demonstrated impressive results in generating high-resolution images of complex and imaginative scenes. Recent approaches have further extended these methods with personalization techniques that…

计算机视觉与模式识别 · 计算机科学 2025-05-05 Tanzila Rahman , Shweta Mahajan , Hsin-Ying Lee , Jian Ren , Sergey Tulyakov , Leonid Sigal

Existing text-based 3D generation methods generate attractive results but lack detailed geometry control. Sketches, known for their conciseness and expressiveness, have contributed to intuitive 3D modeling but are confined to producing…

图形学 · 计算机科学 2024-05-15 Feng-Lin Liu , Hongbo Fu , Yu-Kun Lai , Lin Gao

The field of generative AI has a transformative impact on various areas, including virtual reality, autonomous driving, the metaverse, gaming, and robotics. Among these applications, 3D object generation techniques are of utmost importance.…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Qinghong Sun , Yangguang Li , ZeXiang Liu , Xiaoshui Huang , Fenggang Liu , Xihui Liu , Wanli Ouyang , Jing Shao

Despite recent advancements in text-to-image generation, most existing methods struggle to create images with multiple objects and complex spatial relationships in the 3D world. To tackle this limitation, we introduce a generic AI system,…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Yanbo Ding , Shaobin Zhuang , Kunchang Li , Zhengrong Yue , Yu Qiao , Yali Wang

Traditional video captioning requests a holistic description of the video, yet the detailed descriptions of the specific objects may not be available. Without associating the moving trajectories, these image-based data-driven methods cannot…

计算机视觉与模式识别 · 计算机科学 2020-07-15 Fangyi Zhu , Jenq-Neng Hwang , Zhanyu Ma , Guang Chen , Jun Guo

Domain generalization (DG) attempts to generalize a model trained on single or multiple source domains to the unseen target domain. Benefiting from the success of Visual-and-Language Pre-trained models in recent years, we argue that it is…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Geng Liu , Yuxi Wang

Object detection, scene graph generation and region captioning, which are three scene understanding tasks at different semantic levels, are tied together: scene graphs are generated on top of objects detected in an image with their pairwise…

计算机视觉与模式识别 · 计算机科学 2017-09-18 Yikang Li , Wanli Ouyang , Bolei Zhou , Kun Wang , Xiaogang Wang

The increasing demand for high-quality 3D content creation has motivated the development of automated methods for creating 3D object models from a single image and/or from a text prompt. However, the reconstructed 3D objects using…

计算机视觉与模式识别 · 计算机科学 2023-04-07 Hoigi Seo , Hayeon Kim , Gwanghyun Kim , Se Young Chun

In this paper, we design and train a Generative Image-to-text Transformer, GIT, to unify vision-language tasks such as image/video captioning and question answering. While generative models provide a consistent network architecture between…

计算机视觉与模式识别 · 计算机科学 2022-12-19 Jianfeng Wang , Zhengyuan Yang , Xiaowei Hu , Linjie Li , Kevin Lin , Zhe Gan , Zicheng Liu , Ce Liu , Lijuan Wang

Text-to-3D generation has attracted much attention from the computer vision community. Existing methods mainly optimize a neural field from scratch for each text prompt, relying on heavy and repetitive training cost which impedes their…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Ming Li , Pan Zhou , Jia-Wei Liu , Jussi Keppo , Min Lin , Shuicheng Yan , Xiangyu Xu

Emerging Metaverse applications demand accessible, accurate, and easy-to-use tools for 3D digital human creations in order to depict different cultures and societies as if in the physical world. Recent large-scale vision-language advances…

图形学 · 计算机科学 2023-04-07 Longwen Zhang , Qiwei Qiu , Hongyang Lin , Qixuan Zhang , Cheng Shi , Wei Yang , Ye Shi , Sibei Yang , Lan Xu , Jingyi Yu

Grounded video description (GVD) encourages captioning models to attend to appropriate video regions (e.g., objects) dynamically and generate a description. Such a setting can help explain the decisions of captioning models and prevents the…

计算机视觉与模式识别 · 计算机科学 2021-12-03 Wenqiao Zhang , Xin Eric Wang , Siliang Tang , Haizhou Shi , Haocheng Shi , Jun Xiao , Yueting Zhuang , William Yang Wang

The topic diversity of open-domain videos leads to various vocabularies and linguistic expressions in describing video contents, and therefore, makes the video captioning task even more challenging. In this paper, we propose an unified…

计算机视觉与模式识别 · 计算机科学 2023-02-15 Shizhe Chen , Jia Chen , Qin Jin , Alexander Hauptmann