中文
相关论文

相关论文: Multi-Dimensional Quality Assessment for Text-to-3…

200 篇论文

Recent text-to-image models have improved global realism, but text rendering remains a persistent failure mode: images may look convincing overall, yet local typography often contains malformed glyphs, broken strokes, irregular spacing, and…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Kirill Koltsov , Aleksandr Gushchin , Anastasia Antsiferova , Dmitriy Vatolin

3D Gaussian Splatting (3DGS) has emerged as a promising approach for novel view synthesis, offering real-time rendering with high visual fidelity. However, its substantial storage requirements present significant challenges for practical…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Yuke Xing , Jiarui Wang , Peizhi Niu , Wenjie Huang , Guangtao Zhai , Yiling Xu

With the rapid advancement of 3D visualization, 3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time, high-fidelity rendering. While prior research has emphasized algorithmic performance and visual fidelity, the…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Zhaolin Wan , Yining Diao , Jingqi Xu , Hao Wang , Zhiyang Li , Xiaopeng Fan , Wangmeng Zuo , Debin Zhao

Automatic 3D content creation seeks to replace labor-intensive modeling and scanning pipelines with systems that can synthesize or recover 3D assets directly from text or images. Its applications span video games, virtual reality, robotics,…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Jiahao Li

This paper explores the burgeoning field of 3D content generation within the landscape of Artificial Intelligence Generated Content (AIGC) and large-scale models. It investigates innovative methods like Text-to-3D and Image-to-3D, which…

图形学 · 计算机科学 2024-05-27 Ke Zhao , Andreas Larsen

In the past utilities relied on in-field inspections to identify asset defects. Recently, utilities have started using drone-based inspections to enhance the field-inspection process. We consider a vast repository of drone images, providing…

Text-to-3D generation has achieved significant success by incorporating powerful 2D diffusion models, but insufficient 3D prior knowledge also leads to the inconsistency of 3D geometry. Recently, since large-scale multi-view datasets have…

计算机视觉与模式识别 · 计算机科学 2024-05-03 Junyoung Seo , Susung Hong , Wooseok Jang , Inès Hyeonsu Kim , Minseop Kwak , Doyup Lee , Seungryong Kim

The digital industry demands high-quality, diverse modular 3D assets, especially for user-generated content~(UGC). In this work, we introduce AssetFormer, an autoregressive Transformer-based model designed to generate modular 3D assets from…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Lingting Zhu , Shengju Qian , Haidi Fan , Jiayu Dong , Zhenchao Jin , Siwei Zhou , Gen Dong , Xin Wang , Lequan Yu

Asset management requires accurate 3D models to inform the maintenance, repair, and assessment of buildings, maritime vessels, and other key structures as they age. These downstream applications rely on high-fidelity models produced from…

计算机视觉与模式识别 · 计算机科学 2026-03-19 James L. Gray , Nikolai Goncharov , Alexandre Cardaillac , Ryan Griffiths , Jack Naylor , Donald G. Dansereau

Recent advances in text-to-video (T2V) technology, as demonstrated by models such as Runway Gen-3, Pika, Sora, and Kling, have significantly broadened the applicability and popularity of the technology. This progress has created a growing…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Zelu Qi , Ping Shi , Shuqi Wang , Chaoyang Zhang , Fei Zhao , Zefeng Ying , Da Pan , Xi Yang , Zheqi He , Teng Dai

The recent advancements in Text-to-Video Artificial Intelligence Generated Content (AIGC) have been remarkable. Compared with traditional videos, the assessment of AIGC videos encounters various challenges: visual inconsistency that defy…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Bowen Qu , Xiaoyu Liang , Shangkun Sun , Wei Gao

The rapid advancement of large multimodal models (LMMs) has led to the rapid expansion of artificial intelligence generated videos (AIGVs), which highlights the pressing need for effective video quality assessment (VQA) models designed…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Jiarui Wang , Huiyu Duan , Guangtao Zhai , Juntong Wang , Xiongkuo Min

When humans read a specific text, they often visualize the corresponding images, and we hope that computers can do the same. Text-to-image synthesis (T2I), which focuses on generating high-quality images from textual descriptions, has…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Nonghai Zhang , Hao Tang

Real-world applications often require a large gallery of 3D assets that share a consistent theme. While remarkable advances have been made in general 3D content creation from text or image, synthesizing customized 3D assets following the…

计算机视觉与模式识别 · 计算机科学 2024-05-16 Zhenwei Wang , Tengfei Wang , Gerhard Hancke , Ziwei Liu , Rynson W. H. Lau

While recent advances in neural representations and generative models have revolutionized 3D content creation, the field remains constrained by significant data processing bottlenecks. To address this, we introduce HY3D-Bench, an…

The labor- and experience-intensive creation of 3D assets with physically based rendering (PBR) materials demands an autonomous 3D asset creation pipeline. However, most existing 3D generation methods focus on geometry modeling, either…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Ze-Xin Yin , Jiaxiong Qiu , Liu Liu , Xinjie Wang , Wei Sui , Zhizhong Su , Jian Yang , Jin Xie

The real-world applications of 3D point clouds have been growing rapidly in recent years, but not much effective work has been dedicated to perceptual quality assessment of colored 3D point clouds. In this work, we first build a large 3D…

图像与视频处理 · 电气工程与系统科学 2021-11-11 Honglei Su , Qi Liu , Zhengfang Duanmu , Wentao Liu , Zhou Wang

Due to the lack of large-scale text-3D correspondence data, recent text-to-3D generation works mainly rely on utilizing 2D diffusion models for synthesizing 3D data. Since diffusion-based methods typically require significant optimization…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Bin-Shih Wu , Hong-En Chen , Sheng-Yu Huang , Yu-Chiang Frank Wang

3D content creation plays a vital role in various applications, such as gaming, robotics simulation, and virtual reality. However, the process is labor-intensive and time-consuming, requiring skilled designers to invest considerable effort…

计算机视觉与模式识别 · 计算机科学 2024-05-16 Chenhan Jiang

Text- or image-to-3D generators and 3D scanners can now produce 3D assets with high-quality shapes and textures. These assets typically consist of a single, fused representation, like an implicit neural field, a Gaussian mixture, or a mesh,…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Minghao Chen , Roman Shapovalov , Iro Laina , Tom Monnier , Jianyuan Wang , David Novotny , Andrea Vedaldi