English
Related papers

Related papers: Multi-Dimensional Quality Assessment for Text-to-3…

200 papers

Recently, text-to-image generation has exhibited remarkable advancements, with the ability to produce visually impressive results. In contrast, text-to-3D generation has not yet reached a comparable level of quality. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Yukang Cao , Yan-Pei Cao , Kai Han , Ying Shan , Kwan-Yee K. Wong

A bottleneck in learning to understand articulated 3D objects is the lack of large and diverse datasets. In this paper, we propose to leverage large language models (LLMs) to close this gap and generate articulated assets at scale. We…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Matt Zhou , Ruining Li , Xiaoyang Lyu , Zhaomou Song , Zhening Huang , Chuanxia Zheng , Christian Rupprecht , Andrea Vedaldi , Shangzhe Wu

Text-to-3D generation has achieved remarkable progress in recent years, yet evaluating these methods remains challenging for two reasons: i) Existing benchmarks lack fine-grained evaluation on different prompt categories and evaluation…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Yujie Zhang , Bingyang Cui , Qi Yang , Zhu Li , Yiling Xu

Generating high-quality 3D assets from a given image is highly desirable in various applications such as AR/VR. Recent advances in single-image 3D generation explore feed-forward models that learn to infer the 3D model of an object without…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Yongwei Chen , Tengfei Wang , Tong Wu , Xingang Pan , Kui Jia , Ziwei Liu

Text-to-3D asset generation has achieved significant optimization under the supervision of 2D diffusion priors. However, when dealing with compositional scenes, existing methods encounter several challenges: 1). failure to ensure that…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Qixuan Li , Chao Wang , Zongjin He , Yan Peng

While 3D generation is progressing rapidly, recent work has often focused on obtaining high-resolution assets, leaving user experience and deployability as afterthoughts. We present AssetGen, a 3D generator that focuses instead on these two…

Recently, image-to-3D approaches have achieved significant results with a natural image as input. However, it is not always possible to access these enriched color input samples in practical applications, where only sketches are available.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Wangguandong Zheng , Haifeng Xia , Rui Chen , Ming Shao , Siyu Xia , Zhengming Ding

We present Dual3D, a novel text-to-3D generation framework that generates high-quality 3D assets from texts in only $1$ minute.The key component is a dual-mode multi-view latent diffusion model. Given the noisy multi-view latents, the 2D…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Xinyang Li , Zhangyu Lai , Linning Xu , Jianfei Guo , Liujuan Cao , Shengchuan Zhang , Bo Dai , Rongrong Ji

3D AI-generated content (AIGC) is a passionate field that has significantly accelerated the creation of 3D models in gaming, film, and design. Despite the development of several groundbreaking models that have revolutionized 3D generation,…

In this paper, we propose a 3D asset-referenced diffusion model for image generation, exploring how to integrate 3D assets into image diffusion models. Existing reference-based image generation methods leverage large-scale pretrained…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Hanzhuo Huang , Qingyang Bao , Zekai Gu , Zhongshuo Du , Cheng Lin , Yuan Liu , Sibei Yang

In recent years, AI generative models have made remarkable progress across various domains, including text generation, image generation, and video generation. However, assessing the quality of text-to-video generation is still in its…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Xinli Yue , Jianhui Sun , Han Kong , Liangchao Yao , Tianyi Wang , Lei Li , Fengyun Rao , Jing Lv , Fan Xia , Yuetang Deng , Qian Wang , Lingchen Zhao

Text-to-music (TTM) generation, which converts textual descriptions into audio, opens up innovative avenues for multimedia creation. Achieving high quality and diversity in this process demands extensive, high-quality data, which are often…

Sound · Computer Science 2025-06-18 Chang Li , Ruoyu Wang , Lijuan Liu , Jun Du , Yixuan Sun , Zilu Guo , Zhenrong Zhang , Yuan Jiang , Jianqing Gao , Feng Ma

Text-driven video editing is rapidly advancing, yet its rigorous evaluation remains challenging due to the absence of dedicated video quality assessment (VQA) models capable of discerning the nuances of editing quality. To address this…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Juntong Wang , Jiarui Wang , Huiyu Duan , Guangtao Zhai , Xiongkuo Min

The complexity of SQL and the spatial semantics of PostGIS create barriers for non-experts working with spatial data. Although large language models can translate natural language into SQL, spatial Text-to-SQL is more error-prone than…

Artificial Intelligence · Computer Science 2026-03-31 Ali Khosravi Kazazi , Zhenlong Li , M. Naser Lessani , Guido Cervone

Assessing action quality is both imperative and challenging due to its significant impact on the quality of AI-generated videos, further complicated by the inherently ambiguous nature of actions within AI-generated video (AIGV). Current…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Zijian Chen , Wei Sun , Yuan Tian , Jun Jia , Zicheng Zhang , Jiarui Wang , Ru Huang , Xiongkuo Min , Guangtao Zhai , Wenjun Zhang

Text-to-image (T2I) generative diffusion models have demonstrated outstanding performance in synthesizing diverse, high-quality visuals from text captions. Several layout-to-image models have been developed to control the generation process…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Ahmad Süleyman , Göksel Biricik

Diffusion-based models have recently revolutionized image generation, achieving unprecedented levels of fidelity. However, consistent generation of high-quality images remains challenging partly due to the lack of conditioning mechanisms…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Khaled Abud , Sergey Lavrushkin , Alexey Kirillov , Dmitriy Vatolin

Recent years have witnessed remarkable advances in artificial intelligence generated content(AIGC), with diverse input modalities, e.g., text, image, video, audio and 3D. The 3D is the most close visual modality to real-world 3D environment…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Jian Liu , Xiaoshui Huang , Tianyu Huang , Lu Chen , Yuenan Hou , Shixiang Tang , Ziwei Liu , Wanli Ouyang , Wangmeng Zuo , Junjun Jiang , Xianming Liu

Data augmentation has been established as an efficacious approach to supplement useful information for low-resource datasets. Traditional augmentation techniques such as noise injection and image transformations have been widely used. In…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Yuwei Yin , Jean Kaddour , Xiang Zhang , Yixin Nie , Zhenguang Liu , Lingpeng Kong , Qi Liu

Recent text-to-3D models can render high-quality assets, yet they still stumble on objects with complex attributes. The key obstacles are: (1) existing text-to-3D approaches typically lift text-to-image models to extract semantics via text…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Yiming Qin , Zhu Xu , Yang Liu
‹ Prev 1 3 4 5 6 7 10 Next ›